Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
Researchers have identified a critical limitation in Multimodal Large Language Models (MLLMs), which remain prone to hallucinations in dynamic video scenes due to inadequate spatio-temporal monitoring. To address this, the team introduced STEMO-Bench, a new benchmark featuring human-verified, object-centric facts that evaluate intermediate reasoning by decomposing queries into sub-questions. This approach distinguishes genuine temporal understanding from coincidental correctness, overcoming the limitations of existing benchmarks that rely on single final-answer evaluations. Furthermore, the authors proposed STEMO-Track, a novel object-centric framework designed to explicitly construct and reason over structured object trajectories through chunk-wise state extraction and temporal aggregation. Extensive experiments demonstrate that STEMO-Track significantly reduces hallucinated answers and enhances spatio-temporal reasoning consistency compared to state-of-the-art MLLMs. This study provides a rigorous diagnostic tool and a robust solution for improving video understanding capabilities in AI systems, marking a significant advancement in computer vision and artificial intelligence research.
Wire timeline
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
Researchers have identified a critical limitation in Multimodal Large Language Models (MLLMs), which remain prone to hallucinations in dynamic video scenes due to inadequate spatio-temporal monitoring. To address this, the team introduced STEMO-Bench, a new benchmark featuring human-verified, object-centric facts that evaluate intermediate reasoning by decomposing queries into sub-questions. This approach distinguishes genuine temporal understanding from coincidental correctness, overcoming the limitations of existing benchmarks that rely on single final-answer evaluations. Furthermore, the authors proposed STEMO-Track, a novel object-centric framework designed to explicitly construct and reason over structured object trajectories through chunk-wise state extraction and temporal aggregation. Extensive experiments demonstrate that STEMO-Track significantly reduces hallucinated answers and enhances spatio-temporal reasoning consistency compared to state-of-the-art MLLMs. This study provides a rigorous diagnostic tool and a robust solution for improving video understanding capabilities in AI systems, marking a significant advancement in computer vision and artificial intelligence research.
cs.AI updates on arXiv.org