Study Suggests Benchmarks Underestimate LLM Performance in Hallucination Detection
A new study published on arXiv investigates whether current benchmarks underestimate the performance of Large Language Models (LLMs) in detecting contextual hallucinations, particularly within summarization tasks. Researchers analyzed the QAGS-C and SummEval datasets by comparing original human annotations with predictions from Gemini 2.5 Flash and GPT-5 Mini. To resolve discrepancies between human labels and model judgments, the team implemented a human adjudication process involving two cross-cultural adjudicators for conflicted samples. The results indicated that triple agreement among humans, GPT, and Gemini increased significantly, rising by 6.38% for QAGS-C and 7.62% for SummEval. Model accuracy also improved, with notable gains for both GPT and Gemini. Crucially, adjudicators frequently favored model judgments when explicit reasoning was provided, suggesting that single-pass human annotations may be insufficient for ambiguity-prone tasks. The findings advocate for model-assisted re-evaluation to create more reliable benchmarks, highlighting the potential of LLMs to enhance evaluation standards in AI development.
Wire timeline
Study Suggests Benchmarks Underestimate LLM Performance in Hallucination Detection
A new study published on arXiv investigates whether current benchmarks underestimate the performance of Large Language Models (LLMs) in detecting contextual hallucinations, particularly within summarization tasks. Researchers analyzed the QAGS-C and SummEval datasets by comparing original human annotations with predictions from Gemini 2.5 Flash and GPT-5 Mini. To resolve discrepancies between human labels and model judgments, the team implemented a human adjudication process involving two cross-cultural adjudicators for conflicted samples. The results indicated that triple agreement among humans, GPT, and Gemini increased significantly, rising by 6.38% for QAGS-C and 7.62% for SummEval. Model accuracy also improved, with notable gains for both GPT and Gemini. Crucially, adjudicators frequently favored model judgments when explicit reasoning was provided, suggesting that single-pass human annotations may be insufficient for ambiguity-prone tasks. The findings advocate for model-assisted re-evaluation to create more reliable benchmarks, highlighting the potential of LLMs to enhance evaluation standards in AI development.
cs.AI updates on arXiv.org