Sem-ECE: A Semantic-Sampling Framework for Evaluating LLM Calibration in Open-Ended QA
Researchers have introduced Sem-ECE (Semantic-Sampling Expected Calibration Error), a new framework designed to evaluate the calibration of Large Language Models (LLMs) in open-ended question answering scenarios. Calibration, which measures the alignment between a model's predicted confidence and its actual accuracy, is critical for deploying LLMs in high-stakes fields like medicine and law. Existing evaluation methods often fail in open-ended settings due to reliance on restricted output formats, self-reported verbalized confidence, or ambiguous sampling rules. The proposed Sem-ECE framework samples multiple answers, groups them into semantic classes, and uses frequency as a confidence metric. It includes two estimators, Sem1-ECE and Sem2-ECE, which are proven to be asymptotically unbiased. Experiments across five leading commercial LLMs demonstrate that Sem-ECE outperforms verbalized confidence and existing sampling-based methods. Furthermore, the divergence between the two estimators serves as a diagnostic tool for question difficulty. This work addresses a significant gap in realistic LLM evaluation, offering a robust alternative when internal probabilities are unavailable.
Wire timeline
Sem-ECE: A Semantic-Sampling Framework for Evaluating LLM Calibration in Open-Ended QA
Researchers have introduced Sem-ECE (Semantic-Sampling Expected Calibration Error), a new framework designed to evaluate the calibration of Large Language Models (LLMs) in open-ended question answering scenarios. Calibration, which measures the alignment between a model's predicted confidence and its actual accuracy, is critical for deploying LLMs in high-stakes fields like medicine and law. Existing evaluation methods often fail in open-ended settings due to reliance on restricted output formats, self-reported verbalized confidence, or ambiguous sampling rules. The proposed Sem-ECE framework samples multiple answers, groups them into semantic classes, and uses frequency as a confidence metric. It includes two estimators, Sem1-ECE and Sem2-ECE, which are proven to be asymptotically unbiased. Experiments across five leading commercial LLMs demonstrate that Sem-ECE outperforms verbalized confidence and existing sampling-based methods. Furthermore, the divergence between the two estimators serves as a diagnostic tool for question difficulty. This work addresses a significant gap in realistic LLM evaluation, offering a robust alternative when internal probabilities are unavailable.
cs.AI updates on arXiv.org