Quantum reinforcement learning improves cost and delay in quantum cloud scheduling

Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud Orchestration

Machine LearningArtificial IntelligenceDistributed, Parallel, and Cluster ComputingEmerging Technologies

Summary

Scheduling tasks on a quantum cloud is hard because different quantum computers vary in cost and speed, and standard pricing doesn’t fit well. The authors propose a new method combining quantum circuits with advanced AI techniques to balance cost and delay better than simple rule-based approaches. Their method uses fewer parameters but achieves similar or better performance, making it more efficient. This approach could help manage quantum cloud resources more effectively in the future.

What this means in practice

  • For quantum cloud operators: Schedule quantum tasks more efficiently by balancing execution cost with delay using quantum-enhanced reinforcement learning.
  • For cloud resource managers: Improve orchestration decisions for hybrid systems with heterogeneous resources by using compact quantum AI models that reduce training complexity.

Authors

An N. H. Phan, Dang Van Huynh, Muhammad Usman, Hoa T. Nguyen

Abstract

Quantum cloud computing, delivered through the quantum-as-a-service (QaaS) model, provides access to quantum computing resources. However, applying uniform time-based pricing across fundamentally heterogeneous quantum resources significantly complicates task orchestration, particularly when addressing the tradeoff between execution costs and system performance. While heuristic methods rely on predefined scheduling rules, classical deep reinforcement learning (DRL) models may require more trainable parameters in this setting. Motivated by the potential of parameterised quantum circuits (PQCs) as compact function approximators, we propose QRLQ, a cost-delay-aware quantum cloud scheduling framework integrating PQCs with a dueling double deep Q-network (D3QN) to dynamically account for both cost and delay. Our simulation results show that QRLQ achieves lower mean cost and delay than the heuristic baselines, achieving a 5-11% lower mean cost relative to availability-based and rotation-based heuristics and reducing mean delay by 17% and 82% relative to the strongest and weakest heuristic baselines, respectively, while retaining execution fidelity within 2% of a fidelity-greedy policy. Compared with the classical DRL baseline, QRLQ achieves comparable scheduling performance while using 72% fewer trainable parameters. This work explores the feasibility of using QRL for task orchestration in quantum cloud environments and demonstrates its potential for cost-delay-aware quantum resource management.