Dynamic tensor memory policies show sudden slowdowns and failures

Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization

Machine LearningDistributed, Parallel, and Cluster Computing

Summary

Training deep neural networks needs careful memory management. The authors studied a memory-saving method called Dynamic Tensor Rematerialization and found it can suddenly slow down a lot or even run out of memory depending on tiny changes in the memory budget. They saw these effects using a computer simulator on two types of networks, LSTM and ResNet-32. The paper explains some causes of these sudden changes but notes that these findings are based on simulations, not yet tested in real systems.

What this means in practice

  • For deep learning engineers: Avoid critical memory budget values that cause slowdowns or crashes during DNN training with dynamic tensor rematerialization.
  • For hardware memory managers: Design memory controllers tuned to prevent recursive rematerialization frontiers that cause out-of-memory errors in constrained deep learning workloads.

Tested on simulated data.

Authors

Mahesh Reddy Pagadala

Abstract

We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated re-eviction of the same storages (evictions per storage rise from 1.33 to 8.27 while the set of distinct evicted storages is essentially unchanged: 5,233 vs 5,236, with the two sets overlapping at Jaccard 0.999). On a ResNet-32 trace, a fine budget sweep reveals a deterministic feasibility inversion: the run is feasible at ratio 0.101, infeasible (OOM) across 0.102-0.106, and feasible again from 0.107. We trace the immediate cause of the OOM to a fully pinned recursive rematerialization frontier that exceeds the budget after every evictable tensor has been evicted. Ablations using the DTR authors' own variants implicate the joint size-staleness scoring term in the observed LSTM instability. We argue these are at least two distinct budget-sensitive pathologies rather than one mechanism, and we separate what is demonstrated from what remains hypothesised. All results concern the reference simulator; reproduction in a production runtime is future work. Code, instrumentation, and raw results accompany this preprint.