LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
Researchers have introduced LegalCiteBench, a new benchmark designed to evaluate the reliability of large language models (LLMs) in generating legal citations. As LLMs become increasingly integrated into legal workflows, the risk of fabricated precedents poses significant professional hazards. Existing benchmarks often overlook this specific failure mode in common-law contexts. LegalCiteBench comprises approximately 24,000 evaluation instances derived from 1,000 real U.S. judicial opinions, covering tasks such as citation retrieval, completion, error detection, and case matching. The study evaluated 21 different LLMs in a closed-book setting, revealing that exact citation recovery remains extremely challenging. Even the most advanced models scored below 7 out of 100 on retrieval and completion tasks. Furthermore, misleading answer rates exceeded 94% for most models on retrieval-heavy tasks. The findings indicate that increasing model scale or applying legal-domain pretraining offers limited improvements. Additionally, prompt-based instructions for uncertainty reduced confident fabrication but did not enhance overall citation correctness, highlighting the need for better diagnostic frameworks for authority generation.
Wire timeline
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
Researchers have introduced LegalCiteBench, a new benchmark designed to evaluate the reliability of large language models (LLMs) in generating legal citations. As LLMs become increasingly integrated into legal workflows, the risk of fabricated precedents poses significant professional hazards. Existing benchmarks often overlook this specific failure mode in common-law contexts. LegalCiteBench comprises approximately 24,000 evaluation instances derived from 1,000 real U.S. judicial opinions, covering tasks such as citation retrieval, completion, error detection, and case matching. The study evaluated 21 different LLMs in a closed-book setting, revealing that exact citation recovery remains extremely challenging. Even the most advanced models scored below 7 out of 100 on retrieval and completion tasks. Furthermore, misleading answer rates exceeded 94% for most models on retrieval-heavy tasks. The findings indicate that increasing model scale or applying legal-domain pretraining offers limited improvements. Additionally, prompt-based instructions for uncertainty reduced confident fabrication but did not enhance overall citation correctness, highlighting the need for better diagnostic frameworks for authority generation.
cs.AI updates on arXiv.org