Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
Researchers have introduced Autonomous Preference Optimization (APO), a novel framework designed to address reasoning alignment challenges in multi-modal large language models (MLLMs) within non-stationary environments. The study highlights how diverse reasoning distributions from source models can transmit systematic biases and drift to target models. APO reformulates this issue as a constraint satisfaction problem, treating inter-model divergences as dynamic negative constraints rather than noise. The framework employs a two-stage protocol: supervised bootstrapping to project the target model into the capability union of source models, followed by constraint-aware optimization using a multi-negative Plackett-Luce objective to suppress drifting trajectories. Extensive experiments on chest X-ray interpretation demonstrate that the proposed 7B model achieves superior robustness, outperforming proprietary source models in average accuracy. Additionally, the team released CXR-MAX, a large-scale benchmark containing 170,982 reasoning trajectories from seven MLLMs, to facilitate further research on reasoning alignment under drift. This work represents a significant advancement in enhancing the reliability and consistency of AI reasoning in dynamic, multi-source settings.
Wire timeline
Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
Researchers have introduced Autonomous Preference Optimization (APO), a novel framework designed to address reasoning alignment challenges in multi-modal large language models (MLLMs) within non-stationary environments. The study highlights how diverse reasoning distributions from source models can transmit systematic biases and drift to target models. APO reformulates this issue as a constraint satisfaction problem, treating inter-model divergences as dynamic negative constraints rather than noise. The framework employs a two-stage protocol: supervised bootstrapping to project the target model into the capability union of source models, followed by constraint-aware optimization using a multi-negative Plackett-Luce objective to suppress drifting trajectories. Extensive experiments on chest X-ray interpretation demonstrate that the proposed 7B model achieves superior robustness, outperforming proprietary source models in average accuracy. Additionally, the team released CXR-MAX, a large-scale benchmark containing 170,982 reasoning trajectories from seven MLLMs, to facilitate further research on reasoning alignment under drift. This work represents a significant advancement in enhancing the reliability and consistency of AI reasoning in dynamic, multi-source settings.
cs.AI updates on arXiv.org