Benchmark reveals varied strengths of language models on quantum tasks
QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks
Artificial IntelligenceMachine Learning
Summary
Evaluating how well large language models understand quantum computing, the authors created a test covering many different quantum tasks like building circuits and fixing errors. They found that a model that does well overall might not do well on some specific tasks, showing that these skills vary a lot. The benchmark runs tasks automatically, so no manual checking is needed, and their methods ensure reliable results. They also shared the test and data openly for others to use.
What this means in practice
- •For quantum software developers: Evaluate how different language models perform on specific quantum programming and debugging tasks to select the best model for their projects.
- •For ai tool integrators: Use QC-Stark benchmark scores to robustly compare language models’ quantum computing capabilities for building smarter AI assistants.
Authors
Pranav Gupta
Abstract
We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.