REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
This research paper introduces REAP (Relevance and Execution-Audited Pipeline), an automated system designed to create reliable benchmarks for AI coding agents using real-world production data. Current evaluation methods, such as online A/B testing and public benchmarks, often suffer from slow feedback loops, irreproducibility, or a lack of alignment with actual production workloads. REAP addresses these issues by curating tasks from genuine developer-agent interactions without requiring manual labeling. It employs an automated verification layer that includes LLM-based task classification, agentic test-relevance validation, and multi-run stability checks to ensure benchmark quality and reliability. The authors demonstrate REAP's effectiveness by creating 'Harvest,' a benchmark featuring tasks across multiple programming languages, primarily Hack. Evaluations of five frontier AI models on Harvest revealed solve rates between 42.9% and 58.2%, highlighting significant capability differences. This approach provides fast, reproducible, and high-fidelity evaluation signals, enabling organizations to make informed decisions regarding the deployment of AI coding assistants in large-scale software engineering environments.
Wire timeline
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
This research paper introduces REAP (Relevance and Execution-Audited Pipeline), an automated system designed to create reliable benchmarks for AI coding agents using real-world production data. Current evaluation methods, such as online A/B testing and public benchmarks, often suffer from slow feedback loops, irreproducibility, or a lack of alignment with actual production workloads. REAP addresses these issues by curating tasks from genuine developer-agent interactions without requiring manual labeling. It employs an automated verification layer that includes LLM-based task classification, agentic test-relevance validation, and multi-run stability checks to ensure benchmark quality and reliability. The authors demonstrate REAP's effectiveness by creating 'Harvest,' a benchmark featuring tasks across multiple programming languages, primarily Hack. Evaluations of five frontier AI models on Harvest revealed solve rates between 42.9% and 58.2%, highlighting significant capability differences. This approach provides fast, reproducible, and high-fidelity evaluation signals, enabling organizations to make informed decisions regarding the deployment of AI coding assistants in large-scale software engineering environments.
cs.AI updates on arXiv.org