BenchCAD: A New Benchmark for Industrial CAD Code Generation
Researchers have introduced BenchCAD, a comprehensive benchmark designed to evaluate the industrial readiness of Multimodal Large Language Models (MLLMs) in generating programmatic Computer-Aided Design (CAD) code. Published on arXiv, this study addresses the gap in evaluating whether AI models can produce executable parametric programs from visual or textual inputs that reflect realistic engineering constraints. BenchCAD comprises 17,900 execution-verified CadQuery programs across 106 industrial part families, such as bevel gears and compression springs. The benchmark assesses models through tasks like image-to-code generation and visual question answering. Evaluations of over ten frontier models reveal that while current systems can recover coarse outer geometry, they frequently fail to generate faithful parametric CAD programs. Common errors include missing fine 3D structures, misinterpreting design parameters, and substituting complex operations with simpler patterns. Although fine-tuning improves performance on known data, generalization to unseen part families remains limited. BenchCAD aims to drive improvements in multimodal CAD automation by providing a standardized metric for perception, parametric abstraction, and executable program synthesis in industrial settings.
Wire timeline
BenchCAD: A New Benchmark for Industrial CAD Code Generation
Researchers have introduced BenchCAD, a comprehensive benchmark designed to evaluate the industrial readiness of Multimodal Large Language Models (MLLMs) in generating programmatic Computer-Aided Design (CAD) code. Published on arXiv, this study addresses the gap in evaluating whether AI models can produce executable parametric programs from visual or textual inputs that reflect realistic engineering constraints. BenchCAD comprises 17,900 execution-verified CadQuery programs across 106 industrial part families, such as bevel gears and compression springs. The benchmark assesses models through tasks like image-to-code generation and visual question answering. Evaluations of over ten frontier models reveal that while current systems can recover coarse outer geometry, they frequently fail to generate faithful parametric CAD programs. Common errors include missing fine 3D structures, misinterpreting design parameters, and substituting complex operations with simpler patterns. Although fine-tuning improves performance on known data, generalization to unseen part families remains limited. BenchCAD aims to drive improvements in multimodal CAD automation by providing a standardized metric for perception, parametric abstraction, and executable program synthesis in industrial settings.
cs.AI updates on arXiv.org