Measuring energy and performance trade offs in high performance computing systems

Measuring Sustainability in Multi-Scale High-Performance Computing

Distributed, Parallel, and Cluster Computing

Summary

Balancing performance and energy use in powerful computer systems is tricky. This paper by the authors introduces a way to measure and understand how different computer setups use energy and perform tasks quickly. They tested a mixed set of computing nodes and found how speed, accuracy, and energy use interact. Their work helps create better strategies to manage and run demanding applications like artificial intelligence and quantum computing more efficiently and sustainably.

High Performance Computing (HPC)Computing ContinuumEnergy EfficiencyThroughputLatencyScalabilitySystem UtilizationAI workloadsQuantum ComputingSchedulers

Authors

Carlos J Barrios, Frédéric Le Mouël, Yves Denneulin

Abstract

The transition from traditional High Performance Computing (HPC) to the Computing Continuum emphasizes efficient resource management and sustainable practices across Multi-Scale hybrid architectures. This paper introduces a multidimensional metric framework to characterize these systems and guide deployment strategies for modern workloads. The framework combines Architectural Performance metrics (such as Throughput, Latency, Scalability), System Utilization, and key Sustainability and Accuracy indicators (such as Energy Efficiency and Power Consumption). Using a modular hybrid testbed, experiments reveal complex relationships among metrics, especially the trade-offs between accuracy and energy, and the efficiency of hybrid nodes. The guidelines help identify optimal operating points and lay the groundwork for improving orchestrators and schedulers (e.g., Kubernetes) to assign demanding applications, including AI and Quantum Computing, to suitable system modules, ensuring high performance and sustainability.