Tag-based Few-shot Learning Improves LLM Reliability in Medical Incident Analysis
A new study published on arXiv introduces a tag-based few-shot example selection method to enhance the reliability of Large Language Models (LLMs) in healthcare settings. The research addresses the critical need for accurate generation of clinical insights, specifically background causal factors and preventive measures, from medical incident reports. Using the Japanese Medical Incident Dataset (JMID), which contains 3,884 real-world accident and near-miss reports, the authors compared three selection strategies: random sampling, cosine similarity-based selection, and their proposed tag-based approach. Experiments conducted with GPT-4o and LLaMA 3.3 demonstrated that the tag-based method achieved superior precision and stability. In contrast, similarity-based selection frequently resulted in unintended outputs and triggered safety filters. The findings suggest that leveraging human-interpretable dataset tags for example selection significantly improves the performance and safety of LLMs in high-stakes clinical applications, offering a robust solution for automated medical incident analysis.
Wire timeline
Tag-based Few-shot Learning Improves LLM Reliability in Medical Incident Analysis
A new study published on arXiv introduces a tag-based few-shot example selection method to enhance the reliability of Large Language Models (LLMs) in healthcare settings. The research addresses the critical need for accurate generation of clinical insights, specifically background causal factors and preventive measures, from medical incident reports. Using the Japanese Medical Incident Dataset (JMID), which contains 3,884 real-world accident and near-miss reports, the authors compared three selection strategies: random sampling, cosine similarity-based selection, and their proposed tag-based approach. Experiments conducted with GPT-4o and LLaMA 3.3 demonstrated that the tag-based method achieved superior precision and stability. In contrast, similarity-based selection frequently resulted in unintended outputs and triggered safety filters. The findings suggest that leveraging human-interpretable dataset tags for example selection significantly improves the performance and safety of LLMs in high-stakes clinical applications, offering a robust solution for automated medical incident analysis.
cs.AI updates on arXiv.org