CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics
Researchers have introduced CLR-voyance, a novel framework designed to enhance open-ended reasoning in inpatient clinical decision support. Addressing the limitations of existing evaluations that often rely on closed-form retrieval or unanchored scoring, this approach reformulates clinical reasoning as a Partially Observable Markov Decision Process (POMDP). The system utilizes outcome-grounded and clinician-validated rewards, partitioning patient journeys into visible pasts and oracle-only futures to generate adaptive, verifiable rubrics. These rubrics guide the post-training of models like Qwen3-8B and MedGemma-4B using GRPO and model merging. The resulting CLR-voyance-8B model achieves state-of-the-art performance, scoring 84.91% on the new CLR-POMDP benchmark, surpassing frontier models such as GPT-5 and MedGemma-27B. A large-scale clinician alignment study validated the framework, providing insights into clinical LLM judging. Furthermore, the system has been successfully deployed in a partner public hospital for over six months, assisting in drafting thousands of complex inpatient notes, demonstrating both technical superiority and practical clinical utility.
Wire timeline
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics
Researchers have introduced CLR-voyance, a novel framework designed to enhance open-ended reasoning in inpatient clinical decision support. Addressing the limitations of existing evaluations that often rely on closed-form retrieval or unanchored scoring, this approach reformulates clinical reasoning as a Partially Observable Markov Decision Process (POMDP). The system utilizes outcome-grounded and clinician-validated rewards, partitioning patient journeys into visible pasts and oracle-only futures to generate adaptive, verifiable rubrics. These rubrics guide the post-training of models like Qwen3-8B and MedGemma-4B using GRPO and model merging. The resulting CLR-voyance-8B model achieves state-of-the-art performance, scoring 84.91% on the new CLR-POMDP benchmark, surpassing frontier models such as GPT-5 and MedGemma-27B. A large-scale clinician alignment study validated the framework, providing insights into clinical LLM judging. Furthermore, the system has been successfully deployed in a partner public hospital for over six months, assisting in drafting thousands of complex inpatient notes, demonstrating both technical superiority and practical clinical utility.
cs.AI updates on arXiv.org