Abstract
Heterogeneous Supercloud systems are transforming data analytics by enabling scalable and efficient task distribution across diverse resources. This paper presents computational models and performance estimation techniques tailored for accelerating Dense Cholesky (CH) Decomposition and Singular Value Decomposition (SVD)-based analytics on Heterogeneous Superclouds. Our ReDSEa tool-chain automates mapping, load balancing, scheduling, parallelism, and computation overlap. Implemented on a heterogeneous system with a Huawei Kunpeng 920 ARM CPU and an Ascend 910 AI accelerator, our LLVM compiler tool-chain employs novel performance models for recursive, iterative, and blocked computations, achieving up to 17x speedup for CH and 88x for SVD over fully optimized 48-core CPU implementations.