PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation
Researchers have introduced PDEAgent-Bench, the first publicly available benchmark specifically designed for evaluating partial differential equation (PDE) to-solver code generation. Unlike existing general-purpose code benchmarks that focus on syntactic correctness, PDEAgent-Bench assesses numerical accuracy, computational efficiency, and compatibility with professional finite-element method (FEM) libraries such as DOLFINx, Firedrake, and deal.II. The benchmark comprises 645 instances across six mathematical categories and eleven PDE families, utilizing a staged evaluation framework requiring generated solvers to pass executability, accuracy, and efficiency checks sequentially. Experiments with representative large language models (LLMs) and code agents reveal that while models frequently produce runnable code, their success rates drop significantly when strict accuracy and efficiency constraints are applied. These findings highlight current limitations in AI agents' ability to generate numerically reliable and efficient PDE solvers. PDEAgent-Bench provides a reproducible testbed grounded in practical numerical solving requirements, aiming to advance the development of robust automated scientific computing tools.
Wire timeline
PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation
Researchers have introduced PDEAgent-Bench, the first publicly available benchmark specifically designed for evaluating partial differential equation (PDE) to-solver code generation. Unlike existing general-purpose code benchmarks that focus on syntactic correctness, PDEAgent-Bench assesses numerical accuracy, computational efficiency, and compatibility with professional finite-element method (FEM) libraries such as DOLFINx, Firedrake, and deal.II. The benchmark comprises 645 instances across six mathematical categories and eleven PDE families, utilizing a staged evaluation framework requiring generated solvers to pass executability, accuracy, and efficiency checks sequentially. Experiments with representative large language models (LLMs) and code agents reveal that while models frequently produce runnable code, their success rates drop significantly when strict accuracy and efficiency constraints are applied. These findings highlight current limitations in AI agents' ability to generate numerically reliable and efficient PDE solvers. PDEAgent-Bench provides a reproducible testbed grounded in practical numerical solving requirements, aiming to advance the development of robust automated scientific computing tools.
cs.AI updates on arXiv.org