MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Researchers have introduced MECAT, a new benchmark designed to address the limitations of current evaluation methods for large audio-language models. Despite advancements in open-ended audio understanding, existing models often lack nuanced, human-level comprehension, a gap exacerbated by benchmarks that fail to distinguish between generic and detailed outputs. MECAT utilizes a pipeline integrating specialized expert models with Chain-of-Thought reasoning from large language models to generate multi-perspective, fine-grained captions and open-set question-answering pairs. Complementing this dataset is DATE (Discriminative-Enhanced Audio Text Evaluation), a novel metric that penalizes vague terms while rewarding detailed descriptions by combining semantic similarity with cross-sample discriminability. The study includes a comprehensive evaluation of state-of-the-art audio models, offering insights into their current capabilities and limitations. Developed by a team including researchers from Xiaomi Research, the project aims to enhance the precision of audio understanding tasks. The associated data and code are publicly available on GitHub, facilitating further research and development in the field of artificial intelligence and audio processing.
Wire timeline
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Researchers have introduced MECAT, a new benchmark designed to address the limitations of current evaluation methods for large audio-language models. Despite advancements in open-ended audio understanding, existing models often lack nuanced, human-level comprehension, a gap exacerbated by benchmarks that fail to distinguish between generic and detailed outputs. MECAT utilizes a pipeline integrating specialized expert models with Chain-of-Thought reasoning from large language models to generate multi-perspective, fine-grained captions and open-set question-answering pairs. Complementing this dataset is DATE (Discriminative-Enhanced Audio Text Evaluation), a novel metric that penalizes vague terms while rewarding detailed descriptions by combining semantic similarity with cross-sample discriminability. The study includes a comprehensive evaluation of state-of-the-art audio models, offering insights into their current capabilities and limitations. Developed by a team including researchers from Xiaomi Research, the project aims to enhance the precision of audio understanding tasks. The associated data and code are publicly available on GitHub, facilitating further research and development in the field of artificial intelligence and audio processing.
cs.AI updates on arXiv.org