Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care
A new study published on arXiv demonstrates that interactive Large Language Models (LLMs) significantly enhance diagnostic accuracy in emergency medicine. The research introduces MedSyn, a system allowing physicians to iteratively query an LLM using full clinical records while initially viewing only the chief complaint. Seven physicians, including three seniors and four residents, evaluated 52 MIMIC-IV cases stratified by difficulty. Results indicated a substantial improvement in diagnostic correctness for residents, particularly in hard cases, where accuracy rose from 0.589 to 0.734. Automated metrics confirmed these gains, showing improved any-match accuracy and F1 scores. Dialogue analysis revealed distinct usage strategies: senior physicians employed targeted, hypothesis-driven questions, whereas residents utilized broader queries. Furthermore, the study noted increased concordance across expertise levels. These findings provide empirical evidence that interactive LLM support meaningfully augments diagnostic reasoning and decision-making workflows in high-pressure emergency care environments, addressing previous gaps in evidence regarding live physician-AI interaction.
Wire timeline
Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care
A new study published on arXiv demonstrates that interactive Large Language Models (LLMs) significantly enhance diagnostic accuracy in emergency medicine. The research introduces MedSyn, a system allowing physicians to iteratively query an LLM using full clinical records while initially viewing only the chief complaint. Seven physicians, including three seniors and four residents, evaluated 52 MIMIC-IV cases stratified by difficulty. Results indicated a substantial improvement in diagnostic correctness for residents, particularly in hard cases, where accuracy rose from 0.589 to 0.734. Automated metrics confirmed these gains, showing improved any-match accuracy and F1 scores. Dialogue analysis revealed distinct usage strategies: senior physicians employed targeted, hypothesis-driven questions, whereas residents utilized broader queries. Furthermore, the study noted increased concordance across expertise levels. These findings provide empirical evidence that interactive LLM support meaningfully augments diagnostic reasoning and decision-making workflows in high-pressure emergency care environments, addressing previous gaps in evidence regarding live physician-AI interaction.
cs.AI updates on arXiv.org