AQUA-Bench: Benchmark for Unanswerable Audio Questions
Researchers have introduced AQUA-Bench, a new benchmark designed to evaluate the reliability of audio-aware large language models in handling unanswerable questions. While current models perform well on standard audio question-answering tasks, they often struggle when no reliable answer can be inferred from the audio input. AQUA-Bench addresses this gap by systematically assessing three specific scenarios: Absent Answer Detection, where the correct option is missing; Incompatible Answer Set Detection, involving categorically mismatched choices; and Incompatible Audio Question Detection, where questions are irrelevant or lack sufficient grounding in the audio. The study highlights that existing benchmarks largely overlook these real-world challenges, leading to a blind spot in current audio-language understanding. By providing a rigorous measure of model performance in these edge cases, AQUA-Bench aims to promote the development of more robust and trustworthy audio-language systems. Experimental results indicate that while models excel at finding answers, they face significant difficulties in recognizing when answers do not exist, underscoring the need for improved assessment tools in artificial intelligence research.
Wire timeline
AQUA-Bench: Benchmark for Unanswerable Audio Questions
Researchers have introduced AQUA-Bench, a new benchmark designed to evaluate the reliability of audio-aware large language models in handling unanswerable questions. While current models perform well on standard audio question-answering tasks, they often struggle when no reliable answer can be inferred from the audio input. AQUA-Bench addresses this gap by systematically assessing three specific scenarios: Absent Answer Detection, where the correct option is missing; Incompatible Answer Set Detection, involving categorically mismatched choices; and Incompatible Audio Question Detection, where questions are irrelevant or lack sufficient grounding in the audio. The study highlights that existing benchmarks largely overlook these real-world challenges, leading to a blind spot in current audio-language understanding. By providing a rigorous measure of model performance in these edge cases, AQUA-Bench aims to promote the development of more robust and trustworthy audio-language systems. Experimental results indicate that while models excel at finding answers, they face significant difficulties in recognizing when answers do not exist, underscoring the need for improved assessment tools in artificial intelligence research.
cs.AI updates on arXiv.org