Vision2cad improves accuracy in parametric cad modeling tasks

Vision2CAD: A Visual Agent Harness for Explicit Geometry Referencing and Localization in Parametric CAD Modeling

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Creating detailed CAD models that can be easily adjusted over time requires careful use of geometric references and constraints. The researchers developed Vision2CAD, a system that helps computers understand and use visual information from sketches more precisely by linking it to the CAD software's internal model. They also created a new dataset to help train and test this system. Experiments show Vision2CAD produces more accurate and stable CAD models compared to previous methods, and it maintains relationships between different model parts when changes are made.

What this means in practice

  • For mechanical engineers: Generate more accurate and editable CAD models by improving geometry referencing and constraint management in parametric designs.
  • For industrial design teams: Use visual information to automate parts of CAD modeling workflows, reducing manual adjustments and errors during early design stages.

Authors

Xi Cheng, Chenxi Zhai, Hang Cheng, Mingyu Fan, Pingfa Feng, Long Zeng

Abstract

Generating parametric CAD models requires accurate geometry and stable feature dependencies. Existing methods face challenges in selecting geometric references, interpreting sketch-plane local coordinates, and establishing sketch constraints to projected external geometry. We present Vision2CAD, a visual agent harness that combines vision-language model (VLM) reasoning with deterministic CAD kernel operations. An ID-based interface supports explicit geometry selection, a local-coordinate bridge converts view coordinates into sketch coordinates, and projected-edge localization supports external sketch constraints. These mechanisms establish feature dependencies within the supported modeling operations and constraint types. We also introduce the Geometry Explicit Reference Dataset (GERD), which aligned commands, geometry states and IDs at every modeling step. On GERD-EVL and a DeepCAD test subset, Vision2CAD improves mIoU by 11.1\% and 5.6\% and reduces Chamfer distance by 17.3\% and 41.8\%, respectively. Parameter-editing experiments and ablation studies further proved the preservation of parametric dependencies.