MDGYM: Benchmarking AI Agents on Molecular Simulations
Researchers have introduced MDGYM, a new benchmark designed to evaluate the capability of AI agents in autonomously executing molecular dynamics (MD) simulations, a critical component of modern scientific discovery. The study assesses whether AI can translate physical intuition into correct computational workflows using widely used packages like LAMMPS and GROMACS. The benchmark comprises 169 expert-curated tasks across three difficulty levels. Evaluations of three agentic frameworks—Claude Code, Codex, and OpenHands—combined with four large language models, revealed significant performance gaps. Even the most capable agents solved only 21% of easy-level tasks, with success rates dropping below 10% for harder challenges. Common failure modes included generating physically unstable configurations, fabricating numerical outputs without actual computation, and premature task abandonment. These findings highlight that fluent code generation does not equate to grounded physical reasoning, distinguishing these failures from those seen in general software engineering benchmarks. The results suggest substantial improvements are needed for AI to reliably support autonomous scientific workflows in physics and chemistry.
Wire timeline
MDGYM: Benchmarking AI Agents on Molecular Simulations
Researchers have introduced MDGYM, a new benchmark designed to evaluate the capability of AI agents in autonomously executing molecular dynamics (MD) simulations, a critical component of modern scientific discovery. The study assesses whether AI can translate physical intuition into correct computational workflows using widely used packages like LAMMPS and GROMACS. The benchmark comprises 169 expert-curated tasks across three difficulty levels. Evaluations of three agentic frameworks—Claude Code, Codex, and OpenHands—combined with four large language models, revealed significant performance gaps. Even the most capable agents solved only 21% of easy-level tasks, with success rates dropping below 10% for harder challenges. Common failure modes included generating physically unstable configurations, fabricating numerical outputs without actual computation, and premature task abandonment. These findings highlight that fluent code generation does not equate to grounded physical reasoning, distinguishing these failures from those seen in general software engineering benchmarks. The results suggest substantial improvements are needed for AI to reliably support autonomous scientific workflows in physics and chemistry.
cs.AI updates on arXiv.org