Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
Researchers from the University of Alberta have published a new study addressing the practical limitations of mixture policies in continuous action reinforcement learning. While theoretically offering greater flexibility than unimodal Gaussian policies, mixture policies have been largely absent from state-of-the-art algorithms due to high variance in gradient estimation. The team proposes a novel Marginalized Reparameterization (MRP) estimator to solve this issue, proving it offers lower variance than standard likelihood-ratio approaches. Extensive experiments across Gym MuJoCo, DeepMind Control Suite, and MetaWorld environments demonstrate that MRP mixture policies significantly outperform previous methods and achieve parity with, or occasionally surpass, Gaussian counterparts. This work clarifies the trade-offs involved in policy representation, effectively transforming mixture policies from a theoretical concept into a viable practical tool for enhancing solution quality and entropy robustness in modern reinforcement learning systems like Soft Actor-Critic.
Wire timeline
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
Researchers from the University of Alberta have published a new study addressing the practical limitations of mixture policies in continuous action reinforcement learning. While theoretically offering greater flexibility than unimodal Gaussian policies, mixture policies have been largely absent from state-of-the-art algorithms due to high variance in gradient estimation. The team proposes a novel Marginalized Reparameterization (MRP) estimator to solve this issue, proving it offers lower variance than standard likelihood-ratio approaches. Extensive experiments across Gym MuJoCo, DeepMind Control Suite, and MetaWorld environments demonstrate that MRP mixture policies significantly outperform previous methods and achieve parity with, or occasionally surpass, Gaussian counterparts. This work clarifies the trade-offs involved in policy representation, effectively transforming mixture policies from a theoretical concept into a viable practical tool for enhancing solution quality and entropy robustness in modern reinforcement learning systems like Soft Actor-Critic.
cs.AI updates on arXiv.org