Simulus: A Modular World Model Agent for Sample-Efficient Reinforcement Learning
Researchers have introduced Simulus, a new modular token-based world model agent designed to enhance sample efficiency in reinforcement learning. Addressing the complexity and integration challenges inherent in current world models, Simulus combines four key improvements inspired by the Rainbow DQN framework: a flexible tokenization system for diverse observation and action modalities, intrinsic motivation to reduce epistemic uncertainty, prioritized world model replay, and a regression-as-classification approach for reward prediction. The study demonstrates that these components yield synergistic gains when combined. Simulus achieves state-of-the-art performance across three distinct benchmarks: visual Atari 100K, continuous-control DMC Proprioception 500K, and symbolic Craftax-1M. Notably, the research highlights that intrinsic motivation remains beneficial even under strict interaction budgets, countering concerns about wasted interactions on irrelevant experiences. Ablation studies confirm that each module contributes individually to the overall performance. The authors, including Lior Cohen and colleagues, have made the code and model weights publicly available, marking a significant advancement in planning-free world model agents for efficient learning environments.
Wire timeline
Simulus: A Modular World Model Agent for Sample-Efficient Reinforcement Learning
Researchers have introduced Simulus, a new modular token-based world model agent designed to enhance sample efficiency in reinforcement learning. Addressing the complexity and integration challenges inherent in current world models, Simulus combines four key improvements inspired by the Rainbow DQN framework: a flexible tokenization system for diverse observation and action modalities, intrinsic motivation to reduce epistemic uncertainty, prioritized world model replay, and a regression-as-classification approach for reward prediction. The study demonstrates that these components yield synergistic gains when combined. Simulus achieves state-of-the-art performance across three distinct benchmarks: visual Atari 100K, continuous-control DMC Proprioception 500K, and symbolic Craftax-1M. Notably, the research highlights that intrinsic motivation remains beneficial even under strict interaction budgets, countering concerns about wasted interactions on irrelevant experiences. Ablation studies confirm that each module contributes individually to the overall performance. The authors, including Lior Cohen and colleagues, have made the code and model weights publicly available, marking a significant advancement in planning-free world model agents for efficient learning environments.
cs.AI updates on arXiv.org