GridProbe: Adaptive Test-Time Compute for Long-Video VLMs
Researchers have introduced GridProbe, a novel training-free inference paradigm designed to optimize long-video understanding in Vision-Language Models (VLMs). Traditional methods struggle with the quadratic attention costs of processing thousands of frames, often relying on pre-training signals that fail in complex reasoning tasks. GridProbe addresses this by arranging frames on a grid and using lightweight posterior probes to generate an interpretable importance map. This map drives a Shape-Adaptive Selection rule, dynamically adjusting the frame budget based on question difficulty without requiring answer access. Empirical results demonstrate significant efficiency gains: on Video-MME-v2, it matches baseline accuracy while reducing computational cost by 3.36 times. On LongVideoBench, it achieves higher accuracy with only 35% of the baseline compute. Furthermore, decoupling the selector from the QA model allows smaller selectors to pair with larger models, yielding superior performance at reduced computational loads. This approach enables adaptive test-time compute, offering a scalable solution for resource-intensive video analysis tasks while maintaining high accuracy and interpretability.
Wire timeline
GridProbe: Adaptive Test-Time Compute for Long-Video VLMs
Researchers have introduced GridProbe, a novel training-free inference paradigm designed to optimize long-video understanding in Vision-Language Models (VLMs). Traditional methods struggle with the quadratic attention costs of processing thousands of frames, often relying on pre-training signals that fail in complex reasoning tasks. GridProbe addresses this by arranging frames on a grid and using lightweight posterior probes to generate an interpretable importance map. This map drives a Shape-Adaptive Selection rule, dynamically adjusting the frame budget based on question difficulty without requiring answer access. Empirical results demonstrate significant efficiency gains: on Video-MME-v2, it matches baseline accuracy while reducing computational cost by 3.36 times. On LongVideoBench, it achieves higher accuracy with only 35% of the baseline compute. Furthermore, decoupling the selector from the QA model allows smaller selectors to pair with larger models, yielding superior performance at reduced computational loads. This approach enables adaptive test-time compute, offering a scalable solution for resource-intensive video analysis tasks while maintaining high accuracy and interpretability.
cs.AI updates on arXiv.org