Probing Cross-modal Information Hubs in Audio-Visual LLMs
Researchers from KAIST have released a new study investigating the internal mechanisms of Audio-Visual Large Language Models (AVLLMs). While AVLLMs are powerful architectures for joint reasoning across audio, visual, and textual modalities, their cross-modal information flow remains largely unexplored compared to text-only or vision-language models. This paper analyzes how information derived from one modality is encoded within the token representations of another. The study uncovers two key findings: AVLLMs primarily encode integrated audio-visual information in specific "sink tokens," and a distinct subset of these, termed "cross-modal sink tokens," specializes in storing such data. Leveraging these insights, the authors propose a simple, training-free method to mitigate hallucinations by encouraging the model to rely on integrated cross-modal information within these specialized tokens. The research aims to deepen the understanding of bidirectional interactions between audio and video modalities in AI systems. The associated code has been made publicly available to support further development and verification of these findings in the field of artificial intelligence.
Wire timeline
Probing Cross-modal Information Hubs in Audio-Visual LLMs
Researchers from KAIST have released a new study investigating the internal mechanisms of Audio-Visual Large Language Models (AVLLMs). While AVLLMs are powerful architectures for joint reasoning across audio, visual, and textual modalities, their cross-modal information flow remains largely unexplored compared to text-only or vision-language models. This paper analyzes how information derived from one modality is encoded within the token representations of another. The study uncovers two key findings: AVLLMs primarily encode integrated audio-visual information in specific "sink tokens," and a distinct subset of these, termed "cross-modal sink tokens," specializes in storing such data. Leveraging these insights, the authors propose a simple, training-free method to mitigate hallucinations by encouraging the model to rely on integrated cross-modal information within these specialized tokens. The research aims to deepen the understanding of bidirectional interactions between audio and video modalities in AI systems. The associated code has been made publicly available to support further development and verification of these findings in the field of artificial intelligence.
cs.AI updates on arXiv.org