Interpretable Coreference Resolution Evaluation Using Explicit Semantics
Researchers from the academic community have introduced a novel, semantically-enhanced evaluation framework for coreference resolution, addressing limitations in traditional aggregate metrics like CoNLL-F1. While standard metrics measure structural overlap, they often fail to provide diagnostic insights into specific semantic categories such as people, locations, or events. This new approach overlays Concept and Named Entity Recognition (CNER) onto coreference outputs, assigning semantic labels to nominal mentions and propagating them to entire clusters. This enables the computation of typed scores that evaluate mention extraction and linking capabilities stratified by semantic class. Experiments conducted on datasets including OntoNotes, LitBank, and PreCo demonstrate that the framework uncovers systematic weaknesses obscured by aggregate metrics. Furthermore, the study shows that these diagnostics facilitate the design of targeted, low-cost data augmentation strategies, resulting in measurable out-of-domain improvements. The paper, submitted to arXiv under Computer Science > Computation and Language, offers a method for more interpretable model assessment and actionable improvements in natural language processing systems.
Wire timeline
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
Researchers from the academic community have introduced a novel, semantically-enhanced evaluation framework for coreference resolution, addressing limitations in traditional aggregate metrics like CoNLL-F1. While standard metrics measure structural overlap, they often fail to provide diagnostic insights into specific semantic categories such as people, locations, or events. This new approach overlays Concept and Named Entity Recognition (CNER) onto coreference outputs, assigning semantic labels to nominal mentions and propagating them to entire clusters. This enables the computation of typed scores that evaluate mention extraction and linking capabilities stratified by semantic class. Experiments conducted on datasets including OntoNotes, LitBank, and PreCo demonstrate that the framework uncovers systematic weaknesses obscured by aggregate metrics. Furthermore, the study shows that these diagnostics facilitate the design of targeted, low-cost data augmentation strategies, resulting in measurable out-of-domain improvements. The paper, submitted to arXiv under Computer Science > Computation and Language, offers a method for more interpretable model assessment and actionable improvements in natural language processing systems.
cs.AI updates on arXiv.org