CrackMeBench: A New Benchmark for Evaluating AI Agents in Binary Reverse Engineering
Researchers Isaac David and Arthur Gervais have introduced CrackMeBench, a novel benchmark designed to evaluate the capabilities of language-model agents in classical binary reverse engineering. Unlike existing benchmarks that focus on source-level code repair or broad cybersecurity capture-the-flag tasks, CrackMeBench specifically tests an agent's ability to recover validation logic from symbol-poor executables and produce accepted inputs or key generators. The benchmark utilizes a secure, no-network Linux Docker sandbox with standard reverse-engineering tools, featuring eight public calibration tasks and twelve generated tasks based on C, Rust, and Go templates. In initial evaluations, GPT-5.5 achieved a 92% pass rate on generated tasks, significantly outperforming Claude Opus 4.7 (58%) and Kimi K2 (42%). The study provides detailed metrics including pass rates, execution time, and tool usage, offering a reproducible testbed for measuring progress in autonomous binary analysis. This development marks a significant step in assessing AI proficiency in low-level software security tasks, moving beyond high-level source code reasoning to precise binary interaction and exploitation within controlled educational environments.
Wire timeline
CrackMeBench: A New Benchmark for Evaluating AI Agents in Binary Reverse Engineering
Researchers Isaac David and Arthur Gervais have introduced CrackMeBench, a novel benchmark designed to evaluate the capabilities of language-model agents in classical binary reverse engineering. Unlike existing benchmarks that focus on source-level code repair or broad cybersecurity capture-the-flag tasks, CrackMeBench specifically tests an agent's ability to recover validation logic from symbol-poor executables and produce accepted inputs or key generators. The benchmark utilizes a secure, no-network Linux Docker sandbox with standard reverse-engineering tools, featuring eight public calibration tasks and twelve generated tasks based on C, Rust, and Go templates. In initial evaluations, GPT-5.5 achieved a 92% pass rate on generated tasks, significantly outperforming Claude Opus 4.7 (58%) and Kimi K2 (42%). The study provides detailed metrics including pass rates, execution time, and tool usage, offering a reproducible testbed for measuring progress in autonomous binary analysis. This development marks a significant step in assessing AI proficiency in low-level software security tasks, moving beyond high-level source code reasoning to precise binary interaction and exploitation within controlled educational environments.
cs.AI updates on arXiv.org