CITE: Anytime-Valid Statistical Inference in LLM Self-Consistency
Researchers have introduced the Certification by Intersection-union Testing with E-processes (CITE) algorithm to address challenges in controlling error levels during Large Language Model (LLM) self-consistency reasoning. While LLMs often improve accuracy by sampling multiple outputs, determining when to stop sampling remains difficult, especially with data-dependent rules and unknown answer sets. The CITE algorithm provides anytime-valid certification of a target answer as the unique mode of the model's response distribution. It provably controls false certification rates at prescribed levels under arbitrary stopping conditions without requiring prior knowledge of the answer category set. The study establishes category-set-size-free stopping-time rates and matching minimax lower bounds. Additionally, the framework extends to confidence-weighted voting. Empirical simulations and experiments demonstrate effective error control and improved certification performance, particularly in diffuse-tail settings. This work contributes significant theoretical guarantees to statistical inference in AI, enhancing the reliability of LLM reasoning processes by ensuring robust statistical validation regardless of when the sampling process is halted.
Wire timeline
CITE: Anytime-Valid Statistical Inference in LLM Self-Consistency
Researchers have introduced the Certification by Intersection-union Testing with E-processes (CITE) algorithm to address challenges in controlling error levels during Large Language Model (LLM) self-consistency reasoning. While LLMs often improve accuracy by sampling multiple outputs, determining when to stop sampling remains difficult, especially with data-dependent rules and unknown answer sets. The CITE algorithm provides anytime-valid certification of a target answer as the unique mode of the model's response distribution. It provably controls false certification rates at prescribed levels under arbitrary stopping conditions without requiring prior knowledge of the answer category set. The study establishes category-set-size-free stopping-time rates and matching minimax lower bounds. Additionally, the framework extends to confidence-weighted voting. Empirical simulations and experiments demonstrate effective error control and improved certification performance, particularly in diffuse-tail settings. This work contributes significant theoretical guarantees to statistical inference in AI, enhancing the reliability of LLM reasoning processes by ensuring robust statistical validation regardless of when the sampling process is halted.
cs.AI updates on arXiv.org