DanLing NestedTensor speeds up deep learning with variable-size inputs
DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning
Machine LearningArtificial IntelligencePerformance
Summary
Handling inputs of different sizes in deep learning often wastes computing power when everything is forced into fixed-size formats with padding. The authors designed DanLing NestedTensor, a new way to represent data that keeps track of varying sizes naturally inside the tensor itself. This approach speeds up computation and reduces memory use by avoiding padding and managing complex data layouts more efficiently. It works transparently with common deep learning tools like PyTorch and maintains support for training and inference.
What this means in practice
- •For deep learning engineers: Implement more efficient models handling inputs of varying sizes without manual padding or offset calculations.
- •For natural language processing teams: Accelerate transformer models like BERT by using variable-size sequence representations that reduce memory and runtime.
Authors
Zhiyuan Chen
Abstract
Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure a property of the tensor itself. Packed values carry tensor-backed partitions and logical dimension order, so broadcasting creates ragged axes, feature transformations retain them, and reductions consume them. The same representation carries through autograd and both eager and compiled execution. On an A100, the geometric-mean speedup over same-mode padding is 2.74$\times$ eager and 3.39$\times$ compiled across four BERT scales, and 1.97$\times$ eager across four FCN backbones. A four-block Pairformer-style workload runs 2.40-4.32$\times$ faster than a padded reference using native PyTorch kernels across square length regimes in eager execution, with peak allocation falling from 38.08 to 5.41 GiB on its high-variation batch. The tensor interface lets model code built from its supported operators compose efficient variable-size computation without managing offsets at any call site. Code will be released publicly upon publication.