CLEAR Framework Reveals Noise and Ambiguity Degrade LLM Reliability in Medicine
Researchers have introduced the CLinical Evaluation of Ambiguity and Reliability (CLEAR) framework to address limitations in current medical Large Language Model (LLM) evaluations. Standard benchmarks often rely on simplified, exam-style questions that fail to capture the ambiguity inherent in real-world medical inquiries. The CLEAR framework systematically assesses how decision-space presentation, ambiguity, and uncertainty impact LLM reasoning by perturbing the number of plausible answers, the presence of abstention options, and semantic framing. Evaluating 17 LLMs across three benchmarks, the study reveals that increasing plausible answer options degrades a model's ability to identify correct answers or abstain from incorrect ones. Furthermore, framing abstention as uncertainty admission, such as "I don't know," rather than assertive rejection, intensifies caution deficits and increases incorrect selections. The authors formalize the gap between identifying correct answers and abstaining from incorrect ones as the "humility deficit," which worsens with model scale. These findings highlight that scaling alone does not resolve reliability issues in medical AI, underscoring the need for more robust evaluation methods that account for noise and ambiguity in clinical contexts.
Wire timeline
CLEAR Framework Reveals Noise and Ambiguity Degrade LLM Reliability in Medicine
Researchers have introduced the CLinical Evaluation of Ambiguity and Reliability (CLEAR) framework to address limitations in current medical Large Language Model (LLM) evaluations. Standard benchmarks often rely on simplified, exam-style questions that fail to capture the ambiguity inherent in real-world medical inquiries. The CLEAR framework systematically assesses how decision-space presentation, ambiguity, and uncertainty impact LLM reasoning by perturbing the number of plausible answers, the presence of abstention options, and semantic framing. Evaluating 17 LLMs across three benchmarks, the study reveals that increasing plausible answer options degrades a model's ability to identify correct answers or abstain from incorrect ones. Furthermore, framing abstention as uncertainty admission, such as "I don't know," rather than assertive rejection, intensifies caution deficits and increases incorrect selections. The authors formalize the gap between identifying correct answers and abstaining from incorrect ones as the "humility deficit," which worsens with model scale. These findings highlight that scaling alone does not resolve reliability issues in medical AI, underscoring the need for more robust evaluation methods that account for noise and ambiguity in clinical contexts.
cs.AI updates on arXiv.org