DOSER: Diffusion-based OOD Detection in Offline Reinforcement Learning
Researchers have introduced DOSER, a novel framework designed to address the critical challenge of out-of-distribution (OOD) action overestimation in offline reinforcement learning (RL). Traditional methods often penalize unseen samples uniformly, which can inaccurately identify OOD actions and suppress beneficial exploration. DOSER overcomes these limitations by employing two diffusion models to capture behavior policy and state distribution, utilizing single-step denoising reconstruction error as a reliable OOD indicator. The framework selectively regularizes actions by distinguishing between beneficial and detrimental OOD moves through predicted transition evaluation, thereby suppressing risky actions while encouraging high-potential exploration. Theoretical analysis proves that DOSER is a gamma-contraction with a unique fixed point and bounded value estimates, offering asymptotic performance guarantees. Extensive benchmarks demonstrate that DOSER consistently outperforms prior methods, particularly on suboptimal datasets. This advancement represents a significant step forward in improving the reliability and efficiency of offline RL systems by moving beyond uniform penalization strategies.
Wire timeline
DOSER: Diffusion-based OOD Detection in Offline Reinforcement Learning
Researchers have introduced DOSER, a novel framework designed to address the critical challenge of out-of-distribution (OOD) action overestimation in offline reinforcement learning (RL). Traditional methods often penalize unseen samples uniformly, which can inaccurately identify OOD actions and suppress beneficial exploration. DOSER overcomes these limitations by employing two diffusion models to capture behavior policy and state distribution, utilizing single-step denoising reconstruction error as a reliable OOD indicator. The framework selectively regularizes actions by distinguishing between beneficial and detrimental OOD moves through predicted transition evaluation, thereby suppressing risky actions while encouraging high-potential exploration. Theoretical analysis proves that DOSER is a gamma-contraction with a unique fixed point and bounded value estimates, offering asymptotic performance guarantees. Extensive benchmarks demonstrate that DOSER consistently outperforms prior methods, particularly on suboptimal datasets. This advancement represents a significant step forward in improving the reliability and efficiency of offline RL systems by moving beyond uniform penalization strategies.
cs.AI updates on arXiv.org