Machine Learning in Healthcare Faces Reproducibility Challenges
Vector Institute Faculty Member Marzyeh Ghassemi, along with Andrew L. Beam and Arjun Manrai, authored a viewpoint article for JAMA addressing the critical challenges of reproducibility and replicability in machine learning models within healthcare. The article distinguishes between reproducibility, which involves obtaining identical results using the same data and code, and replicability, which requires consistent outcomes across different datasets. While machine learning offers significant potential for improving patient care and reducing hospital costs through clinical prediction tools, the authors warn that high-capacity models introduce complexity. Technical choices made during installation or by default can significantly alter outcomes, leading to varied results when models are applied in different clinical settings. This variability poses serious risks to patient safety as these technologies are increasingly deployed in hospitals. Ghassemi emphasizes the urgent need for robust results, advocating for a research culture that prioritizes code sharing and data accessibility. The authors argue that without intellectual verification and transparent practices, the medical community risks creating niche, unverifiable results that could compromise clinical decision-making and long-term patient safety.
Wire timeline
Machine Learning in Healthcare Faces Reproducibility Challenges
Vector Institute Faculty Member Marzyeh Ghassemi, along with Andrew L. Beam and Arjun Manrai, authored a viewpoint article for JAMA addressing the critical challenges of reproducibility and replicability in machine learning models within healthcare. The article distinguishes between reproducibility, which involves obtaining identical results using the same data and code, and replicability, which requires consistent outcomes across different datasets. While machine learning offers significant potential for improving patient care and reducing hospital costs through clinical prediction tools, the authors warn that high-capacity models introduce complexity. Technical choices made during installation or by default can significantly alter outcomes, leading to varied results when models are applied in different clinical settings. This variability poses serious risks to patient safety as these technologies are increasingly deployed in hospitals. Ghassemi emphasizes the urgent need for robust results, advocating for a research culture that prioritizes code sharing and data accessibility. The authors argue that without intellectual verification and transparent practices, the medical community risks creating niche, unverifiable results that could compromise clinical decision-making and long-term patient safety.
Vector Institute for Artificial Intelligence