Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Researchers from arXiv have introduced a novel framework called Grounded Correspondence for video object-centric learning, challenging the conventional approach of using learned dynamics modules to maintain temporal consistency. The study demonstrates that existing predictors are often expensive approximations of discrete correspondence problems. By leveraging modern self-supervised vision backbones, which already encode instance-discriminative features, the new method eliminates the need for learned temporal prediction. Instead, it employs deterministic bipartite matching, specifically Hungarian matching, to maintain frame-to-frame identity based on slot representations initialized from salient regions in frozen backbone features. Notably, this approach requires zero learnable parameters for temporal modeling yet achieves competitive performance on standard benchmarks such as MOVi-D, MOVi-E, and YouTube-VIS. This development suggests a significant shift towards more efficient and effective methods in computer vision, reducing computational costs while maintaining accuracy in object tracking and representation across video frames.
Wire timeline
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Researchers from arXiv have introduced a novel framework called Grounded Correspondence for video object-centric learning, challenging the conventional approach of using learned dynamics modules to maintain temporal consistency. The study demonstrates that existing predictors are often expensive approximations of discrete correspondence problems. By leveraging modern self-supervised vision backbones, which already encode instance-discriminative features, the new method eliminates the need for learned temporal prediction. Instead, it employs deterministic bipartite matching, specifically Hungarian matching, to maintain frame-to-frame identity based on slot representations initialized from salient regions in frozen backbone features. Notably, this approach requires zero learnable parameters for temporal modeling yet achieves competitive performance on standard benchmarks such as MOVi-D, MOVi-E, and YouTube-VIS. This development suggests a significant shift towards more efficient and effective methods in computer vision, reducing computational costs while maintaining accuracy in object tracking and representation across video frames.
cs.AI updates on arXiv.org