EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
Researchers have introduced EcoGym, a new open-source benchmark designed to evaluate the long-horizon planning and execution capabilities of Large Language Models (LLMs) within interactive economic environments. Addressing limitations in current episodic and domain-specific frameworks, EcoGym features three diverse scenarios: Vending, Freelance, and Operation. These environments utilize standardized interfaces and budgeted actions over effectively unbounded horizons, exceeding 1000 steps. The benchmark assesses agents based on business-relevant metrics such as net worth, income, and Daily Active Users (DAU), focusing on strategic coherence and robustness under conditions of partial observability and stochasticity. Experiments involving eleven leading LLMs revealed that no single model dominates across all scenarios, highlighting significant suboptimality in either high-level strategy or efficient action execution. Released as an extensible testbed, EcoGym aims to facilitate transparent evaluation of autonomous agents and study controllability-utility trade-offs in economic settings, marking a significant advancement in AI agent assessment methodology.
Wire timeline
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
Researchers have introduced EcoGym, a new open-source benchmark designed to evaluate the long-horizon planning and execution capabilities of Large Language Models (LLMs) within interactive economic environments. Addressing limitations in current episodic and domain-specific frameworks, EcoGym features three diverse scenarios: Vending, Freelance, and Operation. These environments utilize standardized interfaces and budgeted actions over effectively unbounded horizons, exceeding 1000 steps. The benchmark assesses agents based on business-relevant metrics such as net worth, income, and Daily Active Users (DAU), focusing on strategic coherence and robustness under conditions of partial observability and stochasticity. Experiments involving eleven leading LLMs revealed that no single model dominates across all scenarios, highlighting significant suboptimality in either high-level strategy or efficient action execution. Released as an extensible testbed, EcoGym aims to facilitate transparent evaluation of autonomous agents and study controllability-utility trade-offs in economic settings, marking a significant advancement in AI agent assessment methodology.
cs.AI updates on arXiv.org