Measuring energy and performance trade offs in high performance computing systems
Measuring Sustainability in Multi-Scale High-Performance Computing
Distributed, Parallel, and Cluster Computing
Summary
Balancing performance and energy use in powerful computer systems is tricky. This paper by the authors introduces a way to measure and understand how different computer setups use energy and perform tasks quickly. They tested a mixed set of computing nodes and found how speed, accuracy, and energy use interact. Their work helps create better strategies to manage and run demanding applications like artificial intelligence and quantum computing more efficiently and sustainably.
High Performance Computing (HPC)Computing ContinuumEnergy EfficiencyThroughputLatencyScalabilitySystem UtilizationAI workloadsQuantum ComputingSchedulers
Authors
Carlos J Barrios, Frédéric Le Mouël, Yves Denneulin
Abstract
The transition from traditional High Performance Computing (HPC) to the Computing Continuum emphasizes efficient resource management and sustainable practices across Multi-Scale hybrid architectures. This paper introduces a multidimensional metric framework to characterize these systems and guide deployment strategies for modern workloads. The framework combines Architectural Performance metrics (such as Throughput, Latency, Scalability), System Utilization, and key Sustainability and Accuracy indicators (such as Energy Efficiency and Power Consumption). Using a modular hybrid testbed, experiments reveal complex relationships among metrics, especially the trade-offs between accuracy and energy, and the efficiency of hybrid nodes. The guidelines help identify optimal operating points and lay the groundwork for improving orchestrators and schedulers (e.g., Kubernetes) to assign demanding applications, including AI and Quantum Computing, to suitable system modules, ensuring high performance and sustainability.