Vision language model improves analog circuit layout analysis by 73 percent

THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout

Machine Learning

Summary

Understanding complex analog circuit layouts is hard because they contain detailed physical shapes that define how circuits work. The authors created THEIA, a big dataset of layout pictures with question and answer pairs to help computers learn these designs better. They used a special vision-language model fine-tuned for analog circuit layouts and showed it performs much better than general models. This means designers can ask questions about circuit layouts in a more natural and helpful way.

What this means in practice

Authors

Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni

Abstract

The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDSII file represents the industry-standard database containing the ultimate and most accurate source of information of the analog circuit, encapsulating the complex physical geometries and parasitic realities that define tape out performance. This paper proposes THEIA, a novel dataset containing thousands of layout images paired with question-answer conversations, along with a benchmark that employs a fine-tuned vision-language model (VLM) to analyze GDSII files of analog circuits, enabling designers to interact with and query physical layouts as intuitive, meaningful entities. Experimental results using thousands of analog designs across five realistic tasks demonstrate that the proposed fine-tuned VLM outperforms state-of-the-art general-purpose VLMs by a significant margin (up to 73%), highlighting a fundamental gap between general-purpose multimodal reasoning and domain-specific layout understanding.