CoNL: Self-Evolving LLMs via Meta-Evaluation for Non-Verifiable Tasks
Researchers Yuan Sui and Bryan Hooi have introduced CoNL, a novel framework designed to train large language models (LLMs) on non-verifiable tasks such as creative writing, dialogue, and ethical reasoning. Traditional training methods struggle with these tasks due to the lack of ground-truth labels, while existing LLM-as-Judge approaches are limited by the evaluator's own quality and biases. CoNL addresses this by unifying generation, evaluation, and meta-evaluation through multi-agent self-play. In this system, multiple agents sharing the same policy engage in structured conversations to propose, critique, and revise solutions. The core innovation lies in measuring critique quality by its ability to help other agents improve their solutions, thereby earning diagnostic rewards. This mechanism creates explicit supervision for meta-evaluation without requiring external judges or ground truth. Experimental results across various benchmarks demonstrate that CoNL achieves consistent improvements over self-rewarding baselines while maintaining stable training dynamics, offering a scalable solution for enhancing LLM capabilities in subjective domains.
Wire timeline
CoNL: Self-Evolving LLMs via Meta-Evaluation for Non-Verifiable Tasks
Researchers Yuan Sui and Bryan Hooi have introduced CoNL, a novel framework designed to train large language models (LLMs) on non-verifiable tasks such as creative writing, dialogue, and ethical reasoning. Traditional training methods struggle with these tasks due to the lack of ground-truth labels, while existing LLM-as-Judge approaches are limited by the evaluator's own quality and biases. CoNL addresses this by unifying generation, evaluation, and meta-evaluation through multi-agent self-play. In this system, multiple agents sharing the same policy engage in structured conversations to propose, critique, and revise solutions. The core innovation lies in measuring critique quality by its ability to help other agents improve their solutions, thereby earning diagnostic rewards. This mechanism creates explicit supervision for meta-evaluation without requiring external judges or ground truth. Experimental results across various benchmarks demonstrate that CoNL achieves consistent improvements over self-rewarding baselines while maintaining stable training dynamics, offering a scalable solution for enhancing LLM capabilities in subjective domains.
cs.AI updates on arXiv.org