Active Testing of Large Language Models via Approximate Neyman Allocation
Researchers have introduced a novel active testing algorithm designed to optimize the evaluation of large language models (LLMs), particularly for generative tasks. As LLMs scale, the computational and labeling costs for reliable evaluation increase significantly, creating a bottleneck. Existing active testing methods often fail with generative outputs, focusing instead on classification. This new approach leverages semantic entropy from surrogate models to stratify the evaluation pool and applies approximate Neyman allocation based on extracted signals. Tested across various language and multimodal benchmarks, the method demonstrates superior performance compared to baseline techniques. It closely tracks the theoretical Oracle-Neyman ideal, achieving up to a 28% reduction in Mean Squared Error (MSE) over uniform sampling and an average budget savings of 22.9%. This development addresses the recurring cost of evaluation from pre-training to test-time scaling, offering a more efficient framework for assessing model performance without compromising accuracy. The findings are detailed in a new paper submitted to arXiv, highlighting significant advancements in cost-effective AI model assessment.
Wire timeline
Active Testing of Large Language Models via Approximate Neyman Allocation
Researchers have introduced a novel active testing algorithm designed to optimize the evaluation of large language models (LLMs), particularly for generative tasks. As LLMs scale, the computational and labeling costs for reliable evaluation increase significantly, creating a bottleneck. Existing active testing methods often fail with generative outputs, focusing instead on classification. This new approach leverages semantic entropy from surrogate models to stratify the evaluation pool and applies approximate Neyman allocation based on extracted signals. Tested across various language and multimodal benchmarks, the method demonstrates superior performance compared to baseline techniques. It closely tracks the theoretical Oracle-Neyman ideal, achieving up to a 28% reduction in Mean Squared Error (MSE) over uniform sampling and an average budget savings of 22.9%. This development addresses the recurring cost of evaluation from pre-training to test-time scaling, offering a more efficient framework for assessing model performance without compromising accuracy. The findings are detailed in a new paper submitted to arXiv, highlighting significant advancements in cost-effective AI model assessment.
cs.AI updates on arXiv.org