CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation
Researchers have introduced CADBench, a unified benchmark designed to evaluate progress in AI-assisted Computer-Aided Design (CAD) program generation. Addressing the fragmentation in existing evaluations, CADBench comprises 18,000 samples across six benchmark families derived from datasets like DeepCAD and Fusion 360. It supports five input modalities, including clean and noisy meshes and various render types, while utilizing six metrics to assess geometric fidelity, executability, and program compactness. The study benchmarks eleven vision-language systems, generating over 1.4 million CAD programs. Results indicate that specialized mesh-to-CAD models significantly outperform general-purpose code-generating Vision-Language Models (VLMs) under idealized conditions. However, VLMs remain unreliable for accurate reconstruction. The analysis identifies three key failure modes: performance degradation with increased geometric complexity, brittleness of specialized models during modality shifts, and inconsistent model rankings across different metrics. CADBench serves as a critical diagnostic testbed for advancing editable 3D reconstruction and multimodal CAD understanding. The benchmark is publicly available on Hugging Face, providing a standardized framework for future research in this domain.
Wire timeline
CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation
Researchers have introduced CADBench, a unified benchmark designed to evaluate progress in AI-assisted Computer-Aided Design (CAD) program generation. Addressing the fragmentation in existing evaluations, CADBench comprises 18,000 samples across six benchmark families derived from datasets like DeepCAD and Fusion 360. It supports five input modalities, including clean and noisy meshes and various render types, while utilizing six metrics to assess geometric fidelity, executability, and program compactness. The study benchmarks eleven vision-language systems, generating over 1.4 million CAD programs. Results indicate that specialized mesh-to-CAD models significantly outperform general-purpose code-generating Vision-Language Models (VLMs) under idealized conditions. However, VLMs remain unreliable for accurate reconstruction. The analysis identifies three key failure modes: performance degradation with increased geometric complexity, brittleness of specialized models during modality shifts, and inconsistent model rankings across different metrics. CADBench serves as a critical diagnostic testbed for advancing editable 3D reconstruction and multimodal CAD understanding. The benchmark is publicly available on Hugging Face, providing a standardized framework for future research in this domain.
cs.AI updates on arXiv.org