Distributionally Robust Token Optimization in RLHF
Researchers have introduced Distributionally Robust Token Optimization (DRTO), a novel approach designed to enhance the robustness of Large Language Models (LLMs) against small shifts in prompt wording, format, or language. While LLMs often perform well on familiar data, they frequently fail on multi-step reasoning tasks when faced with minor distributional changes. DRTO addresses this vulnerability by combining token-level Reinforcement Learning from Human Feedback (RLHF) with Distributionally Robust Optimization (DRO). The method constructs f-divergence ambiguity sets over span-level actor losses, allowing the model to principledly emphasize difficult response segments during policy optimization. Empirical evaluations demonstrate that DRTO significantly improves consistency under distribution shifts across various reasoning benchmarks. Specifically, the approach achieved a 4.4 percentage point improvement on the MATH-500 benchmark and a 2.7 percentage point gain on LiveCodeBench compared to standard Robust Token Optimization (RTO). This development represents a significant step forward in making LLMs more reliable for complex reasoning tasks where input variations are common.
Wire timeline
Distributionally Robust Token Optimization in RLHF
Researchers have introduced Distributionally Robust Token Optimization (DRTO), a novel approach designed to enhance the robustness of Large Language Models (LLMs) against small shifts in prompt wording, format, or language. While LLMs often perform well on familiar data, they frequently fail on multi-step reasoning tasks when faced with minor distributional changes. DRTO addresses this vulnerability by combining token-level Reinforcement Learning from Human Feedback (RLHF) with Distributionally Robust Optimization (DRO). The method constructs f-divergence ambiguity sets over span-level actor losses, allowing the model to principledly emphasize difficult response segments during policy optimization. Empirical evaluations demonstrate that DRTO significantly improves consistency under distribution shifts across various reasoning benchmarks. Specifically, the approach achieved a 4.4 percentage point improvement on the MATH-500 benchmark and a 2.7 percentage point gain on LiveCodeBench compared to standard Robust Token Optimization (RTO). This development represents a significant step forward in making LLMs more reliable for complex reasoning tasks where input variations are common.
cs.AI updates on arXiv.org