Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs
Researchers have introduced a new framework called Separate First, Fuse Later (SFFL) to address cross-modal interference in audio-visual large language models (AV-LLMs). Current AV-LLMs often suffer from hallucinations because information from one modality, such as audio, can misguide the interpretation of another, like vision, during intermediate reasoning stages. The SFFL framework mitigates this by enforcing modality-specific chain-of-thought reasoning, which generates separate reasoning traces for audio and visual inputs before integrating evidence for the final answer. The team constructed modality-preference labels through a specialized data pipeline and used them as auxiliary rewards in reinforcement learning to encourage instance-dependent preferences for specific modality cues. Additionally, a modality-specific reasoning mechanism ensures isolation during the separated stage while allowing full cross-modal access during evidence fusion. Experimental results demonstrate that SFFL significantly improves both accuracy and robustness, achieving an average relative gain of 5.16% on general audio-visual question answering benchmarks and an 11.17% improvement on benchmarks specifically testing for cross-modal hallucinations.
Wire timeline
Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs
Researchers have introduced a new framework called Separate First, Fuse Later (SFFL) to address cross-modal interference in audio-visual large language models (AV-LLMs). Current AV-LLMs often suffer from hallucinations because information from one modality, such as audio, can misguide the interpretation of another, like vision, during intermediate reasoning stages. The SFFL framework mitigates this by enforcing modality-specific chain-of-thought reasoning, which generates separate reasoning traces for audio and visual inputs before integrating evidence for the final answer. The team constructed modality-preference labels through a specialized data pipeline and used them as auxiliary rewards in reinforcement learning to encourage instance-dependent preferences for specific modality cues. Additionally, a modality-specific reasoning mechanism ensures isolation during the separated stage while allowing full cross-modal access during evidence fusion. Experimental results demonstrate that SFFL significantly improves both accuracy and robustness, achieving an average relative gain of 5.16% on general audio-visual question answering benchmarks and an 11.17% improvement on benchmarks specifically testing for cross-modal hallucinations.
cs.AI updates on arXiv.org