Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
A recent study published on arXiv challenges the foundational assumption of Learning from Human Feedback (LHF) in AI development, specifically within mental health applications. Researchers tested whether aggregated expert judgments provide valid ground truth for training Large Language Models (LLMs). Three certified psychiatrists independently evaluated LLM-generated responses using a calibrated rubric. The results revealed consistently poor inter-rater reliability, with Intraclass Correlation Coefficients ranging from 0.087 to 0.295, well below acceptable thresholds. Disagreement was most pronounced and systematic in high-stakes categories like suicide and self-harm, where one factor showed negative reliability. Qualitative analysis indicated that these divergences stem from coherent but incompatible clinical frameworks—such as safety-first versus engagement-centered approaches—rather than random measurement error. The findings suggest that current consensus-based aggregation methods erase nuanced professional philosophies, creating arithmetic compromises that may undermine AI safety. Consequently, the authors recommend shifting towards alignment methods that preserve and learn from expert disagreement, characterizing the issue as a sociotechnical phenomenon requiring new approaches to reward modeling and evaluation benchmarks in safety-critical AI systems.
Wire timeline
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
A recent study published on arXiv challenges the foundational assumption of Learning from Human Feedback (LHF) in AI development, specifically within mental health applications. Researchers tested whether aggregated expert judgments provide valid ground truth for training Large Language Models (LLMs). Three certified psychiatrists independently evaluated LLM-generated responses using a calibrated rubric. The results revealed consistently poor inter-rater reliability, with Intraclass Correlation Coefficients ranging from 0.087 to 0.295, well below acceptable thresholds. Disagreement was most pronounced and systematic in high-stakes categories like suicide and self-harm, where one factor showed negative reliability. Qualitative analysis indicated that these divergences stem from coherent but incompatible clinical frameworks—such as safety-first versus engagement-centered approaches—rather than random measurement error. The findings suggest that current consensus-based aggregation methods erase nuanced professional philosophies, creating arithmetic compromises that may undermine AI safety. Consequently, the authors recommend shifting towards alignment methods that preserve and learn from expert disagreement, characterizing the issue as a sociotechnical phenomenon requiring new approaches to reward modeling and evaluation benchmarks in safety-critical AI systems.
cs.AI updates on arXiv.org