CAMEL: Confidence-Gated Reflection for Reward Modeling
Researchers have introduced CAMEL, a novel confidence-gated reflection framework designed to enhance reward modeling for aligning large language models with human preferences. Addressing the trade-off between the efficiency of scalar discriminative models and the interpretability of generative judging models, CAMEL utilizes the log-probability margin between verdict tokens as a proxy for instance difficulty. This allows the system to make lightweight single-token preference decisions initially, selectively invoking deeper reflection only for low-confidence instances. The model is trained using reinforcement learning with counterfactual prefix augmentation to encourage effective self-correction. Empirical results demonstrate that CAMEL achieves state-of-the-art performance across three major benchmarks, attaining an average accuracy of 82.9%. This represents a 3.2% improvement over previous best models. Notably, the 14-billion-parameter CAMEL model outperforms existing 70-billion-parameter models, establishing a superior accuracy-efficiency Pareto frontier. This development signifies a significant advancement in optimizing computational resources while maintaining high alignment quality in AI systems.
Wire timeline
CAMEL: Confidence-Gated Reflection for Reward Modeling
Researchers have introduced CAMEL, a novel confidence-gated reflection framework designed to enhance reward modeling for aligning large language models with human preferences. Addressing the trade-off between the efficiency of scalar discriminative models and the interpretability of generative judging models, CAMEL utilizes the log-probability margin between verdict tokens as a proxy for instance difficulty. This allows the system to make lightweight single-token preference decisions initially, selectively invoking deeper reflection only for low-confidence instances. The model is trained using reinforcement learning with counterfactual prefix augmentation to encourage effective self-correction. Empirical results demonstrate that CAMEL achieves state-of-the-art performance across three major benchmarks, attaining an average accuracy of 82.9%. This represents a 3.2% improvement over previous best models. Notably, the 14-billion-parameter CAMEL model outperforms existing 70-billion-parameter models, establishing a superior accuracy-efficiency Pareto frontier. This development signifies a significant advancement in optimizing computational resources while maintaining high alignment quality in AI systems.
cs.AI updates on arXiv.org