TrajViT: Efficient Video Tokenization via Panoptic Sub-object Trajectories
Researchers have introduced TrajViT, a novel video encoder that revolutionizes video tokenization for transformer models by utilizing panoptic sub-object trajectories instead of traditional space-time patches. This approach, termed grounded video tokenization, aligns tokens with scene complexity rather than video duration, significantly reducing computational redundancy while maintaining temporal coherence. In comparative benchmarks, TrajViT outperforms the standard ViT3D model by achieving a 6% higher top-5 recall in video-text retrieval tasks with ten times fewer tokens. Furthermore, when integrated as an encoder for Video Large Language Models (VideoLLMs), it delivers a 5.2% average performance improvement across six VideoQA benchmarks. The model also offers substantial efficiency gains, featuring four times faster training speeds and requiring 18 times fewer inference FLOPs. This development addresses critical scalability issues in processing long videos, offering a robust solution that enhances both accuracy and computational efficiency in video understanding tasks. The findings suggest a significant leap forward in making video analysis more accessible and less resource-intensive for modern AI applications.
Wire timeline
TrajViT: Efficient Video Tokenization via Panoptic Sub-object Trajectories
Researchers have introduced TrajViT, a novel video encoder that revolutionizes video tokenization for transformer models by utilizing panoptic sub-object trajectories instead of traditional space-time patches. This approach, termed grounded video tokenization, aligns tokens with scene complexity rather than video duration, significantly reducing computational redundancy while maintaining temporal coherence. In comparative benchmarks, TrajViT outperforms the standard ViT3D model by achieving a 6% higher top-5 recall in video-text retrieval tasks with ten times fewer tokens. Furthermore, when integrated as an encoder for Video Large Language Models (VideoLLMs), it delivers a 5.2% average performance improvement across six VideoQA benchmarks. The model also offers substantial efficiency gains, featuring four times faster training speeds and requiring 18 times fewer inference FLOPs. This development addresses critical scalability issues in processing long videos, offering a robust solution that enhances both accuracy and computational efficiency in video understanding tasks. The findings suggest a significant leap forward in making video analysis more accessible and less resource-intensive for modern AI applications.
cs.AI updates on arXiv.org