Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation
Researchers have proposed the Compressed Video Aggregator (CVA), a lightweight module designed to enhance micro-video recommendation systems by decoupling video information from preference learning. The CVA aggregates frozen Video Foundation Model (VFM) embeddings and employs latent reasoning without cross-attention projection to generate compact video embeddings. Addressing issues of frame redundancy and coarse sampling in existing benchmarks, the team utilized titles to re-select key frames based on CLIP technology. Experimental results on the MicroLens and Short-Video datasets demonstrate that CVA achieves consistent performance gains while significantly reducing training time and GPU memory usage by orders of magnitude. Furthermore, the study highlights that re-selected frames can improve the performance of various methods, including CVA. The authors also analyzed the impact of erroneous titles on the method's effectiveness. This academic submission, authored by Yang Xiao and colleagues, aims to provide a more efficient and scalable solution for video recommendation tasks. The associated code is scheduled for release soon, offering potential advancements for developers and researchers in machine learning and computer vision fields focusing on efficient media processing.
Wire timeline
Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation
Researchers have proposed the Compressed Video Aggregator (CVA), a lightweight module designed to enhance micro-video recommendation systems by decoupling video information from preference learning. The CVA aggregates frozen Video Foundation Model (VFM) embeddings and employs latent reasoning without cross-attention projection to generate compact video embeddings. Addressing issues of frame redundancy and coarse sampling in existing benchmarks, the team utilized titles to re-select key frames based on CLIP technology. Experimental results on the MicroLens and Short-Video datasets demonstrate that CVA achieves consistent performance gains while significantly reducing training time and GPU memory usage by orders of magnitude. Furthermore, the study highlights that re-selected frames can improve the performance of various methods, including CVA. The authors also analyzed the impact of erroneous titles on the method's effectiveness. This academic submission, authored by Yang Xiao and colleagues, aims to provide a more efficient and scalable solution for video recommendation tasks. The associated code is scheduled for release soon, offering potential advancements for developers and researchers in machine learning and computer vision fields focusing on efficient media processing.
cs.AI updates on arXiv.org