EgoMemReason: A New Benchmark for Long-Horizon Egocentric Video Reasoning
Researchers have introduced EgoMemReason, a comprehensive benchmark designed to evaluate memory-driven reasoning in long-horizon egocentric video understanding. Addressing the limitations of existing benchmarks that focus primarily on perception and recognition, this new tool targets the challenges faced by next-generation visual assistants like smart glasses and embodied agents. These systems must process continuous visual experiences over days, requiring them to accumulate information, recall prior states, and abstract patterns. EgoMemReason assesses three specific memory types: entity memory for tracking object state changes, event memory for ordering activities across time, and behavior memory for identifying recurring patterns. The benchmark includes 500 questions requiring an average of 25.9 hours of memory backtracking. Evaluations of 17 methods across multimodal large language models and agentic frameworks revealed significant performance gaps, with the best model achieving only 39.6% accuracy. The results highlight that long-horizon memory remains a critical unsolved challenge in AI, providing a foundational resource for advancing memory-aware multimodal systems.
Wire timeline
EgoMemReason: A New Benchmark for Long-Horizon Egocentric Video Reasoning
Researchers have introduced EgoMemReason, a comprehensive benchmark designed to evaluate memory-driven reasoning in long-horizon egocentric video understanding. Addressing the limitations of existing benchmarks that focus primarily on perception and recognition, this new tool targets the challenges faced by next-generation visual assistants like smart glasses and embodied agents. These systems must process continuous visual experiences over days, requiring them to accumulate information, recall prior states, and abstract patterns. EgoMemReason assesses three specific memory types: entity memory for tracking object state changes, event memory for ordering activities across time, and behavior memory for identifying recurring patterns. The benchmark includes 500 questions requiring an average of 25.9 hours of memory backtracking. Evaluations of 17 methods across multimodal large language models and agentic frameworks revealed significant performance gaps, with the best model achieving only 39.6% accuracy. The results highlight that long-horizon memory remains a critical unsolved challenge in AI, providing a foundational resource for advancing memory-aware multimodal systems.
cs.AI updates on arXiv.org