EvoPref: Multi-Objective Evolutionary Optimization for Diverse LLM Alignment
Researchers have introduced EvoPref, a novel multi-objective evolutionary algorithm designed to address preference collapse in large language model (LLM) alignment. Unlike traditional gradient-based methods that often converge to narrow behavioral modes, EvoPref utilizes Non-dominated Sorting Genetic Algorithm II (NSGA-II) selection with archive-based diversity preservation. It optimizes populations of Low-Rank Adaptation (LoRA) adapters across helpfulness, harmlessness, and honesty objectives. Empirical results demonstrate that EvoPref significantly outperforms standard baselines like ORPO, improving preference coverage by 18% and reducing collapse rates by 47%, while maintaining competitive alignment quality on RewardBench. The study provides theoretical motivation based on recent multi-objective evolutionary algorithm runtime analysis, suggesting that population-based methods effectively escape local optima. Comprehensive statistical testing against various evolutionary and gradient baselines confirms the efficacy of this approach. This work establishes evolutionary optimization as a principled paradigm for achieving diverse and robust LLM alignments, offering a significant advancement over single-trajectory optimization techniques currently dominant in the field.
Wire timeline
EvoPref: Multi-Objective Evolutionary Optimization for Diverse LLM Alignment
Researchers have introduced EvoPref, a novel multi-objective evolutionary algorithm designed to address preference collapse in large language model (LLM) alignment. Unlike traditional gradient-based methods that often converge to narrow behavioral modes, EvoPref utilizes Non-dominated Sorting Genetic Algorithm II (NSGA-II) selection with archive-based diversity preservation. It optimizes populations of Low-Rank Adaptation (LoRA) adapters across helpfulness, harmlessness, and honesty objectives. Empirical results demonstrate that EvoPref significantly outperforms standard baselines like ORPO, improving preference coverage by 18% and reducing collapse rates by 47%, while maintaining competitive alignment quality on RewardBench. The study provides theoretical motivation based on recent multi-objective evolutionary algorithm runtime analysis, suggesting that population-based methods effectively escape local optima. Comprehensive statistical testing against various evolutionary and gradient baselines confirms the efficacy of this approach. This work establishes evolutionary optimization as a principled paradigm for achieving diverse and robust LLM alignments, offering a significant advancement over single-trajectory optimization techniques currently dominant in the field.
cs.AI updates on arXiv.org