Study Finds AI Fails Primary Diagnosis Over 80% of Time
A new study published in JAMA Network Open reveals that generative artificial intelligence models fail to produce appropriate differential diagnoses more than 80% of the time, indicating they are not yet safe for unsupervised clinical use. Researchers from Mass General Brigham evaluated 21 large language models, including advanced versions of Claude, GPT, Gemini, and Grok, using 29 standardized clinical vignettes. The assessment tool, PrIME-LLM, tested the models' reasoning capabilities by gradually providing patient information, simulating real-world diagnostic processes. While the AI models demonstrated high accuracy in determining final diagnoses when complete data was available, they struggled significantly with the initial, open-ended stage of clinical reasoning known as differential diagnosis. The study highlights that despite improvements in reasoning-optimized models, off-the-shelf AI lacks the necessary judgment for independent medical deployment. Experts emphasize that human oversight remains indispensable, reinforcing the need for a 'human in the loop' in healthcare settings. The findings suggest that while AI is a promising辅助 tool, it cannot yet replicate the complex reasoning required for safe, standalone clinical decision-making.
Wire timeline
Study Finds AI Fails Primary Diagnosis Over 80% of Time
A new study published in JAMA Network Open reveals that generative artificial intelligence models fail to produce appropriate differential diagnoses more than 80% of the time, indicating they are not yet safe for unsupervised clinical use. Researchers from Mass General Brigham evaluated 21 large language models, including advanced versions of Claude, GPT, Gemini, and Grok, using 29 standardized clinical vignettes. The assessment tool, PrIME-LLM, tested the models' reasoning capabilities by gradually providing patient information, simulating real-world diagnostic processes. While the AI models demonstrated high accuracy in determining final diagnoses when complete data was available, they struggled significantly with the initial, open-ended stage of clinical reasoning known as differential diagnosis. The study highlights that despite improvements in reasoning-optimized models, off-the-shelf AI lacks the necessary judgment for independent medical deployment. Experts emphasize that human oversight remains indispensable, reinforcing the need for a 'human in the loop' in healthcare settings. The findings suggest that while AI is a promising辅助 tool, it cannot yet replicate the complex reasoning required for safe, standalone clinical decision-making.
euronews