MAGIC-Video: Structured Memory for Ultra-Long Agentic Video Reasoning
Researchers have introduced MAGIC-Video, a novel training-free framework designed to address the challenges of understanding ultra-long videos, such as egocentric recordings and surveillance footage spanning days or weeks. Current multimodal large language models struggle with these durations due to limited frame budgets and fragmented memory retrieval. MAGIC-Video utilizes a multimodal memory graph with six typed edges to unify episodic, semantic, and visual content, alongside an interleaved narrative chain that distills long-horizon entity biographies and recurring events. During inference, an agentic loop combines graph retrieval with narrative fact injection, effectively covering both modality and time dimensions in a single pipeline. The framework demonstrates significant performance improvements over existing baselines on benchmarks like EgoLifeQA, Ego-R1, and MM-Lifelong, achieving gains of up to 10.1 points. This advancement enables more coherent and comprehensive analysis of long-duration video data without requiring additional model training, marking a significant step forward in computer vision and artificial intelligence applications for continuous video monitoring and analysis.
Wire timeline
MAGIC-Video: Structured Memory for Ultra-Long Agentic Video Reasoning
Researchers have introduced MAGIC-Video, a novel training-free framework designed to address the challenges of understanding ultra-long videos, such as egocentric recordings and surveillance footage spanning days or weeks. Current multimodal large language models struggle with these durations due to limited frame budgets and fragmented memory retrieval. MAGIC-Video utilizes a multimodal memory graph with six typed edges to unify episodic, semantic, and visual content, alongside an interleaved narrative chain that distills long-horizon entity biographies and recurring events. During inference, an agentic loop combines graph retrieval with narrative fact injection, effectively covering both modality and time dimensions in a single pipeline. The framework demonstrates significant performance improvements over existing baselines on benchmarks like EgoLifeQA, Ego-R1, and MM-Lifelong, achieving gains of up to 10.1 points. This advancement enables more coherent and comprehensive analysis of long-duration video data without requiring additional model training, marking a significant step forward in computer vision and artificial intelligence applications for continuous video monitoring and analysis.
cs.AI updates on arXiv.org