Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports
Researchers have introduced a novel method for extracting data from semi-structured clinical reports that have been digitized via Optical Character Recognition (OCR). Addressing the challenge of fragmented patient records across healthcare institutions due to privacy silos, the study formulates the extraction problem as canonical key-conditioned extractive question answering. The team developed a metric called 'key coverage' to quantify the completeness of a canonical key inventory, which is maintained through iterative mining, normalization, and clustering. Experiments conducted on real-world reports from over 20 hospitals demonstrate that model performance improves monotonically with key coverage. Using a lightweight 0.2B BERT-based model, the system achieved F1 scores of 0.839 for exact match and 0.893 for boundary-tolerant matching when covering the top 90 canonical keys. Notably, this small model outperformed a fine-tuned Qwen3-0.6B baseline under exact match conditions. Although the annotated corpus used was Chinese, the method is language-agnostic and adaptable to other settings, offering a low-cost, on-premise solution for improving Electronic Health Record integration and downstream clinical applications.
Wire timeline
Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports
Researchers have introduced a novel method for extracting data from semi-structured clinical reports that have been digitized via Optical Character Recognition (OCR). Addressing the challenge of fragmented patient records across healthcare institutions due to privacy silos, the study formulates the extraction problem as canonical key-conditioned extractive question answering. The team developed a metric called 'key coverage' to quantify the completeness of a canonical key inventory, which is maintained through iterative mining, normalization, and clustering. Experiments conducted on real-world reports from over 20 hospitals demonstrate that model performance improves monotonically with key coverage. Using a lightweight 0.2B BERT-based model, the system achieved F1 scores of 0.839 for exact match and 0.893 for boundary-tolerant matching when covering the top 90 canonical keys. Notably, this small model outperformed a fine-tuned Qwen3-0.6B baseline under exact match conditions. Although the annotated corpus used was Chinese, the method is language-agnostic and adaptable to other settings, offering a low-cost, on-premise solution for improving Electronic Health Record integration and downstream clinical applications.
cs.AI updates on arXiv.org