Virtual neural networks boost model accuracy without extra parameters
Virtual neural networks: hundreds of souls in a body
Computer Vision and Pattern Recognition
Summary
Making computer programs smarter often means making them bigger and more complex, which can be costly. This paper introduces virtual neural networks, a way to create many separate network models that share the same underlying parts but act independently. These virtual networks all look at the same information, and their combined output is more accurate and reliable than bigger single models. The authors show that adding more virtual networks improves the overall results without needing more memory or bigger models.
What this means in practice
- •For machine learning engineers: Build more accurate and robust image recognition systems without increasing model size by using virtual neural network ensembles.
- •For automated decision system developers: Improve reliability of decisions by deploying ensembles of virtual neural networks sharing weights but producing diverse outputs from the same input data.
Authors
Petr Hurtik, Marek Vajgl, Zahra Alijani, Vojtech Molek
Abstract
A new concept, termed virtual neural networks, is introduced, where the count of trainable parameters is kept constant, and scalability is attained purely through computational resources. This concept is an abstract framework that can be realized using any standard convolutional neural network. It merges siamese neural networks with a deep ensemble technique by generating numerous virtual models that share weights derived from a small set of physical models. The ensemble comprises up to hundreds of trained models simultaneously. All virtual networks take the same input, and their interconnected structure induces an internal distortion that boosts the entire ensemble robustness. The accuracy of the ensemble improves as the number of virtual networks increases, without changing the capacity. Virtual neural networks outperform larger capacity models, typical deep ensembles, and contemporary approaches like SWA and Masksembles. Additionally, the highest-performing individual model from the ensemble surpasses other models trained individually, even those with a greater number of parameters. Code: gitlab.com/EnginCZ/virtual-models-public