ReLU neural networks can be disrupted by training data ordering and poisoning
Let the Neurons Die: Exploiting ReLU-Induced Model Degradation
Machine LearningArtificial Intelligence
Summary
Neural networks using ReLU can have neurons that 'die' by always outputting zero, which stops the network from learning well. The authors show that by carefully choosing the order of training examples or adding specially crafted fake samples, they can make these neurons die more often. This degrades the network's accuracy without changing the network itself. Their experiments on handwriting recognition show noticeable drops in test accuracy using these techniques.
What this means in practice
- •For machine learning engineers: Design better training procedures that avoid data orders or poisoned samples causing neuron death and reduced model accuracy.
- •For security analysts: Detect or defend against attacks that exploit ReLU neuron death through training data manipulation to impair model performance.
Tested on one dataset.
Authors
Kexin Li, Wenjun Qiu, Joshua Abraham, Aditi Maheshwari, David Lie
Abstract
Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availability attacks based on data ordering and poisoning. We begin with the basic dynamic data-ordering attack (DOA), which greedily constructs a training prefix by selecting the next example that minimizes the target layer's post-update weight sum, aiming to push ReLU units toward negative pre-activations without modifying training samples or labels. We then develop two poisoning attacks, IG-DOA and IG-SKA, which use gradient inversion to synthesize class-conditioned samples by matching reference gradients in adverse model states constructed through data ordering or soft knockout, respectively. Soft knockout rearranges weights across adjacent layers to concentrate negative contributions. On a fully connected ReLU network trained on MNIST, ordering 100 of 60,000 training examples reduces test accuracy from 96% to 95% after only five epochs. Adding 200 poisoned samples from a single class reduces test accuracy to approximately 86-88% after five epochs in most evaluated conditions, compared with approximately 96% under clean training. These results demonstrate that ReLU-targeted data ordering and poisoning can impair learning without directly modifying the victim model's parameters.