WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records
Researchers have introduced WISTERIA, a novel weakly supervised representation learning framework designed to address the challenges of noisy and heterogeneous labels in Electronic Health Records (EHR). Traditional EHR modeling often treats clinical labels as ground truth, ignoring the inherent noise from billing codes and incomplete annotations. WISTERIA models these labels as stochastic observations of latent clinical states, employing multiple weak supervision operators to enforce consistency across induced label distributions. This multi-view approach creates an implicit denoising mechanism that reconciles disagreements between noisy labelers. Additionally, the framework incorporates ontology-aware regularization to impose semantic structure on supervision signals. Empirical results demonstrate that WISTERIA significantly improves predictive performance on standard EHR benchmarks, exhibits strong robustness to label noise, and achieves superior cross-institutional generalization compared to existing sequence-based pretraining objectives. The study suggests that explicitly modeling the supervision process provides a more effective inductive bias for learning robust, clinically meaningful representations from complex healthcare data.
Wire timeline
WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records
Researchers have introduced WISTERIA, a novel weakly supervised representation learning framework designed to address the challenges of noisy and heterogeneous labels in Electronic Health Records (EHR). Traditional EHR modeling often treats clinical labels as ground truth, ignoring the inherent noise from billing codes and incomplete annotations. WISTERIA models these labels as stochastic observations of latent clinical states, employing multiple weak supervision operators to enforce consistency across induced label distributions. This multi-view approach creates an implicit denoising mechanism that reconciles disagreements between noisy labelers. Additionally, the framework incorporates ontology-aware regularization to impose semantic structure on supervision signals. Empirical results demonstrate that WISTERIA significantly improves predictive performance on standard EHR benchmarks, exhibits strong robustness to label noise, and achieves superior cross-institutional generalization compared to existing sequence-based pretraining objectives. The study suggests that explicitly modeling the supervision process provides a more effective inductive bias for learning robust, clinically meaningful representations from complex healthcare data.
cs.AI updates on arXiv.org