GPU failure warnings improved with fault-specific prediction model

From Noisy Telemetry to Actionable Warnings: GPU Failure Prediction in Industrial Clusters

Software Engineering

Summary

GPU clusters power many AI services but predicting when GPUs will fail is hard. The authors studied real data from a ByteDance GPU cluster and found challenges like noisy signals and diverse failure types. They developed Falcon, a system that learns specific warning patterns for different faults, improving accuracy and giving early alerts hours in advance. Falcon was tested and deployed in production, showing it can reduce false alarms while catching important GPU failures earlier.

What this means in practice

  • For data center operators: Use fault-specific prediction to detect GPU problems hours before failure, improving maintenance scheduling and reducing downtime.
  • For cloud service providers: Implement calibrated alerts from noisy telemetry to reduce costly false positives while maintaining early warnings of GPU faults in large clusters.

Authors

Yongqian Sun, Run Zhu, Wenwei Gu, Mengyao Li, Shenglin Zhang, Guanjin Wang, Yang Zhang, Xin Wu, Linlin Han, Feng Wang, Xiaozhou Liu, Yu Zhang

Abstract

GPU clusters are critical infrastructure for AI services, but accurate and actionable GPU failure prediction remains a problem in production settings. We study ticket-linked telemetry from a ByteDance GPU cluster and identify three obstacles: workload-confounded telemetry, heterogeneous fault precursors, and the gap between window-level predictions and actionable alerts. These findings motivate Falcon, a fault-specific warning framework combining missingness-aware temporal and peer-relative features, fault-specific learner selection, and an event policy based on thresholding, persistence, and cooldown. On the test set, Falcon achieves the highest F1 among four baselines and reaches 70.6% F1 on the best-performing fault type. Detected cases provide median lead times of 17.34-35.57 hours. We further report a production deployment, where Falcon is calibrated toward high-precision alerts to reflect false-positive costs. Together, these results show that fault-specific modeling improves early warning from noisy production GPU telemetry.