Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria
Researchers have introduced Auto-Rubric as Reward (ARR), a novel framework designed to improve the alignment of multimodal generative models with human preferences. Traditional Reinforcement Learning from Human Feedback (RLHF) methods often collapse nuanced, multi-dimensional human judgments into opaque scalar labels, leading to vulnerabilities like reward hacking. ARR addresses this by externalizing a Vision-Language Model's internal preference knowledge into explicit, prompt-specific rubrics before any pairwise comparison occurs. This approach translates holistic intent into independently verifiable quality dimensions, significantly reducing evaluation biases such as positional bias. Additionally, the study proposes Rubric Policy Optimization (RPO), which converts these structured evaluations into robust binary rewards to stabilize policy gradients during training. Benchmarks in text-to-image generation and image editing demonstrate that ARR-RPO outperforms existing pairwise reward models and VLM judges. The findings suggest that the primary bottleneck in multimodal alignment is not a lack of knowledge, but the absence of a factorized interface for expressing implicit preferences, offering a more reliable and data-efficient solution for generative AI development.
Wire timeline
Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria
Researchers have introduced Auto-Rubric as Reward (ARR), a novel framework designed to improve the alignment of multimodal generative models with human preferences. Traditional Reinforcement Learning from Human Feedback (RLHF) methods often collapse nuanced, multi-dimensional human judgments into opaque scalar labels, leading to vulnerabilities like reward hacking. ARR addresses this by externalizing a Vision-Language Model's internal preference knowledge into explicit, prompt-specific rubrics before any pairwise comparison occurs. This approach translates holistic intent into independently verifiable quality dimensions, significantly reducing evaluation biases such as positional bias. Additionally, the study proposes Rubric Policy Optimization (RPO), which converts these structured evaluations into robust binary rewards to stabilize policy gradients during training. Benchmarks in text-to-image generation and image editing demonstrate that ARR-RPO outperforms existing pairwise reward models and VLM judges. The findings suggest that the primary bottleneck in multimodal alignment is not a lack of knowledge, but the absence of a factorized interface for expressing implicit preferences, offering a more reliable and data-efficient solution for generative AI development.
cs.AI updates on arXiv.org