HY-Himmel: Hierarchical Multi-stream Motion Encoding for Efficient Long Video Understanding
Researchers have introduced HY-Himmel, a novel hierarchical video-language framework designed to address critical bottlenecks in long-video understanding using multimodal language models. Current methods struggle with high decoding costs for dense RGB frames, quadratic token growth, and poor motion perception from sparse sampling. HY-Himmel solves these issues by separating semantic and motion capacity. It routes sparse anchor I-frames to a host Vision Transformer for object identity and scene layout, while employing a lightweight compressed-domain tri-stream adapter for inter-frame intervals. This adapter distills motion evidence from motion-vector maps, residual maps, and context into aligned motion tokens. These tokens are injected into the Large Language Model via a differentiable placeholder mechanism after contrastive alignment. Benchmark results on Video-MME show that HY-Himmel outperforms the dense 32-frame baseline by 2.3 percentage points, achieving 63.5% accuracy while utilizing 3.6 times fewer context tokens. Extensive ablation studies confirm the necessity of the full tri-stream composition for these performance gains, marking a significant advancement in efficient video processing technology.
Wire timeline
HY-Himmel: Hierarchical Multi-stream Motion Encoding for Efficient Long Video Understanding
Researchers have introduced HY-Himmel, a novel hierarchical video-language framework designed to address critical bottlenecks in long-video understanding using multimodal language models. Current methods struggle with high decoding costs for dense RGB frames, quadratic token growth, and poor motion perception from sparse sampling. HY-Himmel solves these issues by separating semantic and motion capacity. It routes sparse anchor I-frames to a host Vision Transformer for object identity and scene layout, while employing a lightweight compressed-domain tri-stream adapter for inter-frame intervals. This adapter distills motion evidence from motion-vector maps, residual maps, and context into aligned motion tokens. These tokens are injected into the Large Language Model via a differentiable placeholder mechanism after contrastive alignment. Benchmark results on Video-MME show that HY-Himmel outperforms the dense 32-frame baseline by 2.3 percentage points, achieving 63.5% accuracy while utilizing 3.6 times fewer context tokens. Extensive ablation studies confirm the necessity of the full tri-stream composition for these performance gains, marking a significant advancement in efficient video processing technology.
cs.AI updates on arXiv.org