Quantum ai agents vary widely in reliable experiment control
Evaluating Verified Autonomy in Quantum Engineering
Artificial Intelligence
Summary
Controlling and testing quantum devices is complex and usually needs a lot of human work. The authors built a virtual lab called Quantum-Harbor where AI agents can try to run quantum experiments safely. They also created a set of 49 tasks called QIQCBench to see how well these agents perform across different challenges. They found that current AI systems show very different levels of reliability, highlighting that being able to do tasks doesn't mean doing them consistently right. This work helps measure how close AI is to reliably running quantum experiments on its own.
What this means in practice
- •For quantum device engineers: Validate autonomous AI tools before using them to run complex quantum calibration and control tasks on real devices.
- •For quantum computing platform operators: Benchmark different AI systems to select reliable agents for automating error correction and system monitoring in quantum processors.
Authors
Naixu Guo, Changhao Li, Siyu Cheng, Qicheng Tang, Binzhao Luo, Bikun Li, Yuxuan Du, Shihao Ru, Jiaqi Cai
Abstract
Reliable quantum engineering is essential for turning quantum phenomena into practical technologies. As quantum platforms grow in scale and complexity, their characterization and operation require increasing human effort and coordination. Scientific artificial intelligence agents, which can plan experiments, operate instruments, and analyze observations, offer a promising route towards autonomous quantum engineering. Yet whether current agents can perform reliably in this setting has not been systematically established. To fill this gap, we developed Quantum-Harbor, a virtual laboratory that provides a controlled execution environment for agents to interact with quantum systems. This design enables direct verification of both the actions taken and the conclusions drawn. Building on this framework, we introduce QIQCBench, a benchmark of $49$ expert-authored tasks spanning multiple layers including calibration and control, error correction and compilation, sensing and networking. Across $17$ frontier agentic systems, QIQCBench reveals wide variation in verified performance. These results expose a substantial gap between demonstrating capability and achieving reliable operation, and establish Quantum-Harbor as a foundation for measuring progress towards verified autonomy in quantum engineering.