VIGIL Framework Disentangles World Completion from Self-Termination in Embodied Agents
Researchers have introduced VIGIL, a new evaluation framework designed to address limitations in standard embodied agent assessments. Current benchmarks often conflate distinct failure modes, such as failing to complete a task, completing it but failing to stop, or reporting success without evidence, into a single metric. VIGIL isolates 'terminal commitment,' the ability of an agent to correctly commit to task completion. Under this protocol, agents rely solely on egocentric RGB inputs without action-success signals and must provide a semantic report verified against the hidden world state. This approach yields two separate scores: world-state completion (W) and benchmark success (B). Testing across 20 models on 1,000 episodes revealed that systems with similar execution capabilities can differ significantly in their ability to report success accurately, with up to a 19.7 percentage point gap in benchmark success. The study highlights that while action-feedback improves execution, it does not necessarily resolve commitment failures, demonstrating the need for independent measurement of terminal commitment in AI development.
Wire timeline
VIGIL Framework Disentangles World Completion from Self-Termination in Embodied Agents
Researchers have introduced VIGIL, a new evaluation framework designed to address limitations in standard embodied agent assessments. Current benchmarks often conflate distinct failure modes, such as failing to complete a task, completing it but failing to stop, or reporting success without evidence, into a single metric. VIGIL isolates 'terminal commitment,' the ability of an agent to correctly commit to task completion. Under this protocol, agents rely solely on egocentric RGB inputs without action-success signals and must provide a semantic report verified against the hidden world state. This approach yields two separate scores: world-state completion (W) and benchmark success (B). Testing across 20 models on 1,000 episodes revealed that systems with similar execution capabilities can differ significantly in their ability to report success accurately, with up to a 19.7 percentage point gap in benchmark success. The study highlights that while action-feedback improves execution, it does not necessarily resolve commitment failures, demonstrating the need for independent measurement of terminal commitment in AI development.
cs.AI updates on arXiv.org