EvoStreaming: Adapting Offline Video Models for Real-Time Streaming Assistance
Researchers have introduced EvoStreaming, a novel framework designed to transform offline Video-Language Models (VideoLLMs) into effective real-time streaming assistants. Current VideoLLMs struggle with the timing of responses in streaming contexts, often lacking policies to balance responsiveness and verbosity. To address this, the team developed RealStreamEval, a new evaluation protocol that penalizes unnecessary responses during sequential observations. The proposed EvoStreaming framework employs a self-evolved adaptation method where the base model generates its own training data, annotates relevance, and determines roll-out policies without external supervision. Remarkably, this approach requires only 1,000 self-generated samples, significantly fewer than leading instruction-tuning methods. Testing across five open VideoLLM backbones, including Qwen and InternVL variants, demonstrated consistent improvements in RealStreamEval scores by up to 10.8 points while preserving offline performance. This study highlights data-efficient interaction tuning as a practical solution for enhancing video understanding models in dynamic, real-time environments.
Wire timeline
EvoStreaming: Adapting Offline Video Models for Real-Time Streaming Assistance
Researchers have introduced EvoStreaming, a novel framework designed to transform offline Video-Language Models (VideoLLMs) into effective real-time streaming assistants. Current VideoLLMs struggle with the timing of responses in streaming contexts, often lacking policies to balance responsiveness and verbosity. To address this, the team developed RealStreamEval, a new evaluation protocol that penalizes unnecessary responses during sequential observations. The proposed EvoStreaming framework employs a self-evolved adaptation method where the base model generates its own training data, annotates relevance, and determines roll-out policies without external supervision. Remarkably, this approach requires only 1,000 self-generated samples, significantly fewer than leading instruction-tuning methods. Testing across five open VideoLLM backbones, including Qwen and InternVL variants, demonstrated consistent improvements in RealStreamEval scores by up to 10.8 points while preserving offline performance. This study highlights data-efficient interaction tuning as a practical solution for enhancing video understanding models in dynamic, real-time environments.
cs.AI updates on arXiv.org