RewardHarness: Self-Evolving Agentic Post-Training Framework for Image Editing Evaluation
Researchers have introduced RewardHarness, a novel self-evolving agentic reward framework designed to evaluate instruction-guided image edits. Addressing the data-efficiency gap in current reward models, which typically require large-scale preference annotations, RewardHarness reframes reward modeling as context evolution rather than weight optimization. The system utilizes an Orchestrator to select relevant tools and skills from a library, while a frozen Sub-Agent constructs reasoning chains to produce preference judgments based on as few as 100 demonstrations. By analyzing successes and failures against ground-truth preferences, the Orchestrator automatically refines its tool library without additional human annotation. Experimental results demonstrate that RewardHarness achieves 47.4% average accuracy on image-editing benchmarks using only 0.05% of standard preference data, surpassing GPT-5 by 5.3 points. Furthermore, when employed as a reward signal for GRPO fine-tuning, it enables RL-tuned models to achieve a score of 3.52 on ImgEdit-Bench. This approach significantly reduces dependency on massive datasets while enhancing alignment with subtle human preferences in AI-generated image editing tasks.
Wire timeline
RewardHarness: Self-Evolving Agentic Post-Training Framework for Image Editing Evaluation
Researchers have introduced RewardHarness, a novel self-evolving agentic reward framework designed to evaluate instruction-guided image edits. Addressing the data-efficiency gap in current reward models, which typically require large-scale preference annotations, RewardHarness reframes reward modeling as context evolution rather than weight optimization. The system utilizes an Orchestrator to select relevant tools and skills from a library, while a frozen Sub-Agent constructs reasoning chains to produce preference judgments based on as few as 100 demonstrations. By analyzing successes and failures against ground-truth preferences, the Orchestrator automatically refines its tool library without additional human annotation. Experimental results demonstrate that RewardHarness achieves 47.4% average accuracy on image-editing benchmarks using only 0.05% of standard preference data, surpassing GPT-5 by 5.3 points. Furthermore, when employed as a reward signal for GRPO fine-tuning, it enables RL-tuned models to achieve a score of 3.52 on ImgEdit-Bench. This approach significantly reduces dependency on massive datasets while enhancing alignment with subtle human preferences in AI-generated image editing tasks.
cs.AI updates on arXiv.org