BUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases
Researchers have introduced BUILD-AND-FIND, a new evaluation protocol designed to assess the quality of codebases generated by AI agents in repository-level engineering contexts. Unlike traditional benchmarks that focus solely on behavioral correctness, this protocol evaluates how effectively downstream agents can recover intended design choices and behaviors from generated code. The method involves a 'builder' agent creating a codebase from a hidden specification and a 'finder' agent attempting to answer specification-traced questions based only on the code. Key metrics include recovery accuracy, repeatability, implementation coverage, and inspection effort. The study emphasizes that lower inspection effort indicates clearer communication of intent within the code artifact. Controls are used to quantify generic priors and separate omitted claims from finder failures. Initial results from a high-prior task pack show near-saturation in recovery accuracy, highlighting inspection effort and finder-specific effects as primary differentiators. This research addresses the growing need for AI-generated code to serve as a clear communication artifact for future automated auditing and extension tasks.
Wire timeline
BUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases
Researchers have introduced BUILD-AND-FIND, a new evaluation protocol designed to assess the quality of codebases generated by AI agents in repository-level engineering contexts. Unlike traditional benchmarks that focus solely on behavioral correctness, this protocol evaluates how effectively downstream agents can recover intended design choices and behaviors from generated code. The method involves a 'builder' agent creating a codebase from a hidden specification and a 'finder' agent attempting to answer specification-traced questions based only on the code. Key metrics include recovery accuracy, repeatability, implementation coverage, and inspection effort. The study emphasizes that lower inspection effort indicates clearer communication of intent within the code artifact. Controls are used to quantify generic priors and separate omitted claims from finder failures. Initial results from a high-prior task pack show near-saturation in recovery accuracy, highlighting inspection effort and finder-specific effects as primary differentiators. This research addresses the growing need for AI-generated code to serve as a clear communication artifact for future automated auditing and extension tasks.
cs.AI updates on arXiv.org