Papers for

deep learning engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Direct feedback alignment reveals common error collapse slows learning

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Abstract: Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitivity across 48 settings. On MNIST, class decodability largely survives collapse, but readout learning remains slow at a fixed learning rate. Adam learns faster despite deeper collapse. Calibrating the baseline readout to the class prior suppresses collapse and speeds learning; weaker feedback trades less collapse for slower learning. Replacing errors by their signs sustains collapse; subtracting the signal's batch mean prevents sustained collapse and improves learning in the tested setting. Related effects occur in deeper and convolutional networks and on CIFAR-10, with severity and cost depending on the readout, optimizer and input statistics.

Fri 25 SeptMachine LearningNeural and Evolutionary Computing
The gist
Training certain neural networks can get stuck because they share a common error pattern that makes hidden units saturate and stop learning effectively. The authors found that this happens when the network’s hidden units respond too similarly, a problem they call 'common-mode collapse.' They show that adjusting the way the output error is fed back and how learning rates are set can reduce this problem and speed up training. Their study tested different settings, networks, and datasets to understand when and why this collapse happens.
Open → 2609.31589v1

Neural network weights optimized to use smaller variable grammar codes

Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights

Abstract: We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.

Fri 25 SeptMachine Learning
The gist
Deep learning models use many numbers called weights, which can take up a lot of space. The authors show a way to tweak these weights so they can be described by shorter, repeating patterns called grammars, saving memory. They do this by grouping similar weight values together and training the model to work well with those simplified weights. This makes the weight representation smaller but with only a small drop in accuracy.
Open → 2609.31564v1

Dynamic tensor memory policies show sudden slowdowns and failures

Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization

Abstract: We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated re-eviction of the same storages (evictions per storage rise from 1.33 to 8.27 while the set of distinct evicted storages is essentially unchanged: 5,233 vs 5,236, with the two sets overlapping at Jaccard 0.999). On a ResNet-32 trace, a fine budget sweep reveals a deterministic feasibility inversion: the run is feasible at ratio 0.101, infeasible (OOM) across 0.102-0.106, and feasible again from 0.107. We trace the immediate cause of the OOM to a fully pinned recursive rematerialization frontier that exceeds the budget after every evictable tensor has been evicted. Ablations using the DTR authors' own variants implicate the joint size-staleness scoring term in the observed LSTM instability. We argue these are at least two distinct budget-sensitive pathologies rather than one mechanism, and we separate what is demonstrated from what remains hypothesised. All results concern the reference simulator; reproduction in a production runtime is future work. Code, instrumentation, and raw results accompany this preprint.

Fri 25 SeptMachine LearningDistributed, Parallel, and Cluster Computing
The gist
Training deep neural networks needs careful memory management. The authors studied a memory-saving method called Dynamic Tensor Rematerialization and found it can suddenly slow down a lot or even run out of memory depending on tiny changes in the memory budget. They saw these effects using a computer simulator on two types of networks, LSTM and ResNet-32. The paper explains some causes of these sudden changes but notes that these findings are based on simulations, not yet tested in real systems.
Open → 2609.31250v1

DanLing NestedTensor speeds up deep learning with variable-size inputs

DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning

Abstract: Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure a property of the tensor itself. Packed values carry tensor-backed partitions and logical dimension order, so broadcasting creates ragged axes, feature transformations retain them, and reductions consume them. The same representation carries through autograd and both eager and compiled execution. On an A100, the geometric-mean speedup over same-mode padding is 2.74$\times$ eager and 3.39$\times$ compiled across four BERT scales, and 1.97$\times$ eager across four FCN backbones. A four-block Pairformer-style workload runs 2.40-4.32$\times$ faster than a padded reference using native PyTorch kernels across square length regimes in eager execution, with peak allocation falling from 38.08 to 5.41 GiB on its high-variation batch. The tensor interface lets model code built from its supported operators compose efficient variable-size computation without managing offsets at any call site. Code will be released publicly upon publication.

Thu 24 SeptMachine LearningArtificial IntelligencePerformance
The gist
Handling inputs of different sizes in deep learning often wastes computing power when everything is forced into fixed-size formats with padding. The authors designed DanLing NestedTensor, a new way to represent data that keeps track of varying sizes naturally inside the tensor itself. This approach speeds up computation and reduces memory use by avoiding padding and managing complex data layouts more efficiently. It works transparently with common deep learning tools like PyTorch and maintains support for training and inference.
Open → 2609.30379v1

KernelOPT improves GPU kernel speed by optimizing compiler outputs

KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization

Abstract: Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kernels, yet treat compiled models as black boxes, generally optimizing individual standalone kernels without respecting the compiler's structural decisions or verifying the model end-to-end. We present KernelOPT, a multi-agent system that treats compiled models as structured artifacts. It preserves vendor library calls (cuBLAS, cuDNN) and exclusively targets generated Triton sub-kernels using five profiling-guided LLM agents. A four-gate verification cascade of static validation, multi-seed correctness, model-level float64-fallback verification, and performance gating filters candidates during optimization and verifies the re-stitched model end-to-end. If no candidate passes all four gates, the system preserves the compiler baseline. The system accepts PyTorch nn.Modules, standalone Triton kernels, and Helion kernels. Evaluated on 250 KernelBench problems, KernelOPT achieves geometric mean speedups over \texttt{torch.compile} of 1.40$\times$ (Level 1: 51/100), 1.15$\times$ (Level 2: 31/100), and 1.07$\times$ (Level 3: 12/50) across all problems.

Thu 24 SeptDistributed, Parallel, and Cluster ComputingArtificial IntelligenceMachine Learning
The gist
Deep learning run times strongly depend on how efficiently GPU code is written, but automatically generated GPU programs are often slower than expert ones. The authors present KernelOPT, a new tool that safely improves GPU code by focusing on specific parts while keeping trusted library calls intact. It uses multiple AI helpers to propose improvements and checks that the overall program still works well before adopting changes. Tests on many problems show KernelOPT can make GPU code run up to 40% faster than current compiler outputs.
Open → 2609.30059v1

Error supervising neural network detects CNN parameter faults

ESupNNet: An Error Supervising Neural Network architecture for error detection against soft errors in parameters

Abstract: This work presents a novel approach to detect misclassification errors in CNNs caused by soft errors in their parameters. We propose an architecture that uses inter-class relations induced by the CNN that needs protection. The architecture has minimal resources overhead and does not require modifying the CNN, which makes it a competent solution that can be used with other error protection techniques. We have validated the architecture with five different combinations of modern dataset-model pairs: ImageNet-1K for ResNet-50 and EfficientNetV2-Small; CIFAR-10 for MobileNetV3, ShuffleNetV2-Small with 2.0x output channels and MNASNet with depth multiplier of 1.3. The validation process was done rigorously with statistical significance, from the creation of the datasets used by the architecture to the acquisition of experimental results. Results show great performance with over 90% accuracy in detecting single errors and great error detection over multiple Bit Error Rates, which can be potentially increased with hyperparameter tuning.

Tue 22 SeptHardware Architecture
The gist
Deep learning models like CNNs can sometimes make mistakes because of tiny errors in their settings caused by things like hardware glitches. The authors created a small extra neural network that watches the main CNN’s behavior to catch when these errors might cause a wrong answer. This helper network uses relationships between categories the CNN knows to spot problems without changing the original model. They tested this system on several popular image recognition models and datasets and found it detects errors with over 90% accuracy.
Open → 2609.26374v1

Neural spectral capacity predicts and optimizes model design efficiently

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

Abstract: Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure: two architectures with identical parameter budgets but different depth-width, head, or FFN allocations receive identical scores yet behave differently. We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix. Under standard random initialization, the Marchenko-Pastur law renders NSC computable from the architectural specification alone, with no model instantiation, data, or gradients. Its layer-wise additive structure admits NSC-DP, an exact dynamic-programming solver returning the architecture globally maximizing NSC under resource constraints in seconds on a CPU -- a guarantee that black-box search over existing training-free proxies cannot provide. Empirically, NSC outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families (on FlexiBERT, $τ= 0.505$ on pairs differing in #Params by less than 10%, where #Params collapses to 0.082); NSC-DP discovers a Transformer-XL architecture on WikiText-103 that beats the human-designed baseline in 2 seconds; and prunes LLaMA-7B to the best 5.7B model across eight commonsense reasoning tasks without any calibration data, about 5900x faster than the strongest training-free proxy baseline.

Sat 19 SeptMachine LearningArtificial Intelligence
The gist
Choosing the right structure for AI models often involves balancing size and computing power, but common measurements miss key details about design. The authors introduce Neural Spectral Capacity (NSC), a quick way to predict a model's potential by looking at its weight matrices without needing to train it or use data. They also provide NSC-DP, a method that quickly finds the best architecture layout under limits like size or compute. Their method ranks models more accurately than traditional measures and speeds up tasks like pruning large models without extra data.
Open → 2609.23087v1

Unified 16-bit format boosts neural network error protection efficiency

UniCASE: A Unified 16-bit Floating-Point Format with Criticality-Aware Selective ECC for Efficient DNN Protection

Abstract: Soft errors are an increasing reliability concern for Deep Neural Network execution because they can corrupt parameters, leading to accuracy degradation. While conventional ECC offers strong fault protection, it incurs additional parity storage and computational overhead. Embedded-parity formats reduce storage cost by reusing the least-significant bits, but they do not optimize protection while reducing computational overhead. We propose UniCASE, a unified 16-bit floating-point (FP) format that jointly optimizes data representation and error protection for reliable DNN execution. It identifies stable blocks across FP64, FP32, FP16, and BFloat16 that can be mapped into a unified representation. Based on bit-level criticality analysis, UniCASE uses selective ECC that assigns distinct levels of protection to different data bits according to their resilience against soft errors. Experimental results show that UniCASE reduces encoder/decoder cost by up to 30%, preserves model accuracy within 1% of the FP32 baseline, and provides significantly stronger soft error resilience than existing embedded-parity methods.

Fri 18 SeptHardware ArchitectureCryptography and Security
The gist
Soft errors can cause mistakes in how deep learning models work, which harms their accuracy. The authors propose UniCASE, a new way to store numbers that both saves space and protects against these errors more efficiently. Their method looks at which parts of the number are most important and protects them better, reducing the overhead of error correction. Tests show UniCASE keeps accuracy close to standard formats while cutting error correction costs and improving reliability.
Open → 2609.22590v1

Matrix adaptive optimization improves neural network training stability

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

Abstract: Adaptive optimization methods such as AdaGrad and Adam are widely used in modern neural-network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop a general Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled approach to deriving matrix-aware adaptive optimization through online regret minimization. By introducing row-wise and column-wise matrix proximal functions and analyzing the resulting regret trade-off, we derive Row-wise Matrix AdaGrad (Row-AdaGrad) and Column-wise Matrix AdaGrad (Column-AdaGrad), with adaptive scaling determined by the accumulated row-wise or column-wise gradient norms. We establish regret guarantees and show that these matrix-aware bounds can be strictly tighter than those of entry-wise AdaGrad under structured gradients. Experiments on matrix factorization and deep neural-network training further demonstrate the benefits of aligning adaptive scaling with matrix structure, including improved optimization stability and trainability at larger learning rates and greater network depths.

Fri 18 SeptMachine LearningArtificial Intelligence
The gist
Optimizing neural networks often involves adjusting many parameters arranged as matrices. Typical methods adapt learning rates for each parameter individually, missing patterns in the matrix structure. The authors develop a new approach that adapts learning rates by looking at entire rows or columns of these matrices, making training more stable and effective. Their method shows better performance when parameters have structured gradients and helps train deeper or larger neural networks.
Open → 2609.21815v1

Tensor program bugs found faster by checking outputs one location at a time

The Output-Space Hypothesis: Enumerative Equivalence Checking for Tensor Programs

Abstract: Tensor programs, as used in deep learning models, are a prime target for optimization, as small performance improvements can have a large impact across training or inference workloads. However, such optimizations are complicated and can produce subtle bugs. Traditionally, correctness is assumed when differential testing against a reference on random inputs fails to reveal bugs. However, the inputs to these programs are massive tensors, and finding bugs can require generating extremely low likelihood inputs with precise relationships among their values. We propose a novel way to find bugs more consistently by flipping the quantifiers. Rather than generating a single input and checking all output tensor locations for equivalence, what if you could check a single output tensor location's equivalence for all inputs? We implement this idea in a system, \dirigo, by using a novel symbolic execution strategy. We demonstrate that \dirigo can find bugs effectively in a public dataset of 6,988 AI-written CUDA kernels that are all marked correct by differential testing. Of these, \dirigo finds 600 kernels that are actually buggy, and finds 97.3\% of those bugs within two minutes.

Thu 17 SeptProgramming LanguagesMachine LearningSoftware Engineering
The gist
Tensor programs used in AI models are tricky to optimize because small bugs can cause big problems. Traditional tests check many outputs for one input, but this misses rare bugs hidden in complex data. The authors flipped this by checking one output location for all possible inputs, catching subtle mistakes more reliably. Their system, Dirigo, found hundreds of bugs that older tests missed, most within minutes.
Open → 2609.19611v1

Sharding technique speeds up transformer retraining and distillation

Affinity-Aware Sharding for Delayed Tensor Parallelism

Abstract: Delayed Tensor Parallelism (DTP) removes the blocking all-reduce of tensor-parallel Transformer inference. Every device adds its own partial output to its residual stream (and broadcasts it) immediately, but only gathers (receives) the other devices' partials $δ$ modules later. A TP to DTP change therefore amounts to a real architecture change, and dense Transformer models need to be retrained or distilled after adaptation. We show that DTP breaks the permutation symmetry of neurons inside FFNs and of KV heads inside attention modules, and that this symmetry breakage makes the sharding itself a modelling decision. We show that maximising the affinity between the KV heads and the FFN neurons co-located on a device, by permuting the dense model before sharding, speeds up the distillation or retraining process. The affinity is measured with a first-order approximation of the damage that losing a head's contribution does to each neuron's output, and the co-located affinity is maximised with a coordinate-ascent optimiser that alternates an exact balanced assignment of neurons with an exhaustive search over the KV head partitions. The whole procedure takes under two minutes on one GPU for Qwen3-0.6B and Danube3-500M. On these models, at $δ=1$, the affinity-optimised layouts reach any distillation target in about half to two thirds of the steps needed by the naive contiguous layouts, over the whole 10k-step range we tested, and every optimised seed beats every contiguous seed and all but one of the sixteen random layouts. We also show that the co-located affinity score at initialisation predicts the KL to the base model after training, across seventeen layouts ranging from anti-optimised to optimised (Pearson $-0.81$ and $-0.89$).

Sat 12 SeptMachine LearningComputation and LanguageDistributed, Parallel, and Cluster Computing
The gist
Transformer models are large AI programs that can be split across many devices to work faster. One way to do this, called Delayed Tensor Parallelism (DTP), changes how parts of the model communicate, which requires retraining the model after splitting it differently. The authors found that rearranging the parts shared on each device to match how they interact speeds up this retraining. They show a method to find the best arrangement of these parts, making the retraining process about twice as fast on tested models.
Open → 2609.13846v1

Lightweight method to check GPU code errors in deep learning training

PEAT: Pseudo-Error Assessment for GPU Kernel Validation in DNN Training

Abstract: Deep neural networks (DNNs) are widely adopted in various fields, driving an emerging trend in developing software stacks associated with DNN training systems. For example, many codes have been ported across different frameworks or developed to leverage the computing power of GPUs or domain-specific accelerators. However, validating a kernel implementation in DNN training is time-consuming and generally requires massive storage. Specifically, this poses a fundamental question: how to characterize the behavior of a new implementation when it is integrated into a DNN training flow. Unfortunately, this problem is not well investigated in the literature, to the best of our knowledge. To address this shortcoming, we present PEAT - a lightweight inspection framework for \underline{P}seudo-\underline{E}rror \underline{A}ssessment associated with GPU kernel validation in DNN \underline{T}raining. Firstly, inspired by conventional fault injection (FI), PEAT's Profiler invokes an operation-wise kernel in a training flow to collect a DNN model's states (e.g., checkpoints and activations). More importantly, the Profiler introduces two simple yet effective techniques, playback FI and frequency-based runtime FI, leveraging persistent kernel calling during the training process. Secondly, PEAT's Analyzer characterizes profiled errors, revealing some signatures from the error distribution of a kernel compared to the golden one. Lastly, PEAT's Detector provides some guidelines as a sufficient condition, which enables associating several well-known error models with signature patterns. We demonstrate the applicability of our approach by presenting the results and analysis using GPUs from the two most popular vendors, NVIDIA V100 and AMD MI250, on various AI models, from vision tasks to language models, for both pretraining and finetuning scenarios.

Fri 11 SeptDistributed, Parallel, and Cluster ComputingComputational Engineering, Finance, and ScienceSoftware Engineering
The gist
Running deep learning programs on GPUs needs special code called kernels, and making sure these kernels work correctly is hard and takes a lot of time and storage. The authors designed a tool named PEAT that watches GPU kernels closely during training, collects information about how the model's data changes, and looks for patterns in errors. This helps identify if the kernels produce unusual errors compared to trusted versions, which means programmers can find problems faster. They tested PEAT with GPU hardware from popular makers and on several AI tasks like image and language processing.
Open → 2609.13544v1

Learning rate schedules inspired by fastest descent improve deep network training

BrachistoneLR: A Brachistochrone-Inspired Learning-Rate Schedule and a Controlled Benchmark of Scheduling Policies

Abstract: The learning-rate schedule is a consequential choice in training deep networks, yet the policies in common use are heuristic, and published comparisons are hard to read, because architecture, dataset, and budget tend to vary alongside the schedule. We study BrachistoneLR, a schedule built by mapping the vertical coordinate of the brachistochrone, the curve of fastest descent under gravity, onto the range between a peak and a floor rate. Expanding the definition shows it to be cosine annealing with the half-period set to E - 1 instead of E, the configuration a standard implementation gives when its period argument is one less than the number of epochs. The rate therefore reaches its floor at the last epoch trained rather than one epoch later, and we show this difference decays as E^-2, making it a short-horizon effect. We then benchmark six schedules over 72 runs on three image classification datasets (MNIST, Fashion-MNIST, CIFAR-10) and four architecture families (fully connected, convolutional, recurrent, residual), fixing the optimizer, data pipeline, and evaluation protocol so that only the schedule varies. Schedules that fall smoothly from peak to floor beat the constant rate and calendar-based decay by margins that grow with task difficulty, reaching 2.5 points of dataset mean on CIFAR-10. Within that leading group, BrachistoneLR, cosine annealing, and warmup-cosine lie within 0.06 accuracy points and 0.17 of a mean rank, which one seed per configuration cannot separate. BrachistoneLR is best on both residual networks and has the highest CIFAR-10 mean, and it sets no milestones, decay factor, warmup length, or restart period. We conclude that the shape of a schedule matters more than its parameterization, that the choice of whether to use a smooth schedule matters more than the choice among them, and that the terminal-rate distinction is worth attention only over short horizons.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Training deep learning models needs careful adjustment of the learning rate, which affects how quickly the model learns. The authors study a new schedule called BrachistoneLR, inspired by the curve of fastest descent in physics, and show it is a slight variation of the popular cosine annealing schedule. They compare six schedules on several datasets and network types while controlling other variables. The results show that smoothly decreasing schedules outperform constant or step-based ones, and BrachistoneLR performs slightly better for some networks without needing extra tuning.
Open → 2609.08069v1

Kan training scales efficiently on multi gpu high performance computers

Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems

Abstract: Kolmogorov-Arnold Networks (KANs) replace the fixed activation functions and linear weights of Multi-Layer Perceptrons (MLPs) with learnable univariate functions on network edges, offering improved interpretability and, in some settings, competitive parameter efficiency. While the approximation properties of KANs have received considerable attention, their behavior under distributed, multi-GPU training has not been systematically characterized. This paper presents an empirical scalability study of data-parallel KAN training on multi-node, multi-GPU high-performance computing (HPC) infrastructure, evaluated along four dimensions: strong scaling, weak scaling, communication overhead, and model-size scaling. Experiments were conducted on the FinisTerrae III supercomputer using up to 8 NVIDIA A100 GPUs across 4 nodes with PyTorch Distributed Data Parallel (DDP). KAN training reaches 74.7% parallel efficiency at 8 GPUs with a 5.97x speedup, consistent with conventional deep learning workloads. Weak scaling shows an initial single-to-multi-GPU throughput drop followed by strong stability. Communication overhead follows a non-monotonic pattern (1.3%-6.1%), driven primarily by All-Reduce algorithm selection and inter-node latency rather than KAN's edge-wise gradient structure. The parameter-to-memory ratio improves with model size even as training time scales unfavorably. These results indicate that operator-level and data-parallel optimizations for KAN are complementary. We provide deployment guidelines for GPU topology and model-size selection, and discuss the limitations of a synthetic-regression evaluation.

Mon 7 SeptDistributed, Parallel, and Cluster ComputingMachine LearningPerformance
The gist
Kolmogorov-Arnold Networks (KANs) are a special type of neural network that use learnable functions instead of fixed math rules to make decisions. The authors studied how well KANs train when the work is split across many GPUs in a supercomputer. They found that KAN training gets faster almost as expected when adding more GPUs, with some small delays due to communication between machines. Their experiments help show how to best set up hardware and model size for training KANs efficiently.
Open → 2609.07740v1

Kolmogorov Arnold stability holds for discontinuous deep learning functions

Kolmogorov--Arnold stability for discontinuous functions

Abstract: Here we investigate the stability of the Kolmogorov--Arnold representation theorem (KART) under adversarial reparameterisations of the hidden layer for multivariate discontinuous and unbounded functions. Our results provide a rigorous mathematical foundation for the structural robustness of modern deep learning architectures, such as Kolmogorov--Arnold Networks (KANs), under adversarial configurations.

Mon 7 SeptMachine Learning
The gist
The paper studies how a mathematical way to represent complex functions, called the Kolmogorov–Arnold theorem, stays reliable even when the function is jumpy or unsteady. The authors explore how this stability works when deep learning networks face tricky changes that try to confuse or break them. Their findings support why certain deep neural networks remain robust when their inner layers are reconfigured in adversarial ways. This helps explain why these networks can perform well even under challenging conditions.
Open → 2609.07240v1