MedMeta: New Benchmark Reveals LLM Limitations in Medical Evidence Synthesis
Researchers have introduced MedMeta, the first benchmark designed to evaluate Large Language Models' (LLMs) ability to synthesize conclusions from medical meta-analyses using study abstracts. Addressing the saturation of standard factual recall benchmarks, MedMeta comprises 81 meta-analyses from PubMed (2018–2025) and assesses models via Retrieval-Augmented Generation (Golden-RAG) and parametric-only workflows. The study validates its LLM-as-a-judge protocol against human expert ratings, showing strong correlation. Results indicate that Golden-RAG significantly outperforms parametric approaches, while domain-specific fine-tuning offers marginal benefits when external data is provided. However, current LLMs achieve only slightly above-average performance (~2.7/5.0) and universally fail to identify negated evidence, highlighting critical vulnerabilities in existing systems. The findings suggest that developing robust RAG systems is more promising for clinical applications than model specialization alone, underscoring the need for better information grounding in AI-driven medical evidence synthesis.
Wire timeline
MedMeta: New Benchmark Reveals LLM Limitations in Medical Evidence Synthesis
Researchers have introduced MedMeta, the first benchmark designed to evaluate Large Language Models' (LLMs) ability to synthesize conclusions from medical meta-analyses using study abstracts. Addressing the saturation of standard factual recall benchmarks, MedMeta comprises 81 meta-analyses from PubMed (2018–2025) and assesses models via Retrieval-Augmented Generation (Golden-RAG) and parametric-only workflows. The study validates its LLM-as-a-judge protocol against human expert ratings, showing strong correlation. Results indicate that Golden-RAG significantly outperforms parametric approaches, while domain-specific fine-tuning offers marginal benefits when external data is provided. However, current LLMs achieve only slightly above-average performance (~2.7/5.0) and universally fail to identify negated evidence, highlighting critical vulnerabilities in existing systems. The findings suggest that developing robust RAG systems is more promising for clinical applications than model specialization alone, underscoring the need for better information grounding in AI-driven medical evidence synthesis.
cs.AI updates on arXiv.org