EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents
Researchers have introduced EnactToM, a new evolving benchmark designed to evaluate functional Theory of Mind (ToM) in embodied AI agents. Unlike existing benchmarks that test literal ToM through direct belief questions, EnactToM assesses the ability of agents to act optimally on implicit beliefs within complex, multi-agent environments. The benchmark features 300 tasks set in a 3D household simulation characterized by partial observability, private information, and constrained communication. Each task is formally verified for solvability and epistemic depth, with difficulty scaling as models improve. Evaluation of seven frontier AI models revealed a significant performance gap: while agents averaged 45.0% on literal belief probes, they scored 0.0% on functional task completion in the hard split. Manual analysis attributed 93% of failures to epistemic coordination breakdowns, such as withheld information and ignored partner constraints. This study highlights critical deficiencies in current AI collaborative capabilities and provides a concrete target for future research in multi-agent systems and artificial intelligence.
Wire timeline
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents
Researchers have introduced EnactToM, a new evolving benchmark designed to evaluate functional Theory of Mind (ToM) in embodied AI agents. Unlike existing benchmarks that test literal ToM through direct belief questions, EnactToM assesses the ability of agents to act optimally on implicit beliefs within complex, multi-agent environments. The benchmark features 300 tasks set in a 3D household simulation characterized by partial observability, private information, and constrained communication. Each task is formally verified for solvability and epistemic depth, with difficulty scaling as models improve. Evaluation of seven frontier AI models revealed a significant performance gap: while agents averaged 45.0% on literal belief probes, they scored 0.0% on functional task completion in the hard split. Manual analysis attributed 93% of failures to epistemic coordination breakdowns, such as withheld information and ignored partner constraints. This study highlights critical deficiencies in current AI collaborative capabilities and provides a concrete target for future research in multi-agent systems and artificial intelligence.
cs.AI updates on arXiv.org