Manifold-Constrained Hyper-Connections for Parameter-Efficient Finetuning
2026-07-20 • Machine Learning
Machine Learning
AI summaryⓘ
The authors explore a new way to tune large language models efficiently by focusing on the residual connections, which are usually left unchanged. They introduce a method called Manifold-Constrained Hyper-Connections (mHC) that adds learnable modules around frozen model parts. They find that mHC alone doesn't always beat an existing method called LoRA, but combining both works better for some tasks and model sizes. Their work suggests that changing residual connections offers a useful new direction for fine-tuning big models.
parameter-efficient finetuningresidual connectionsTransformersManifold-Constrained Hyper-ConnectionsLoRAfrozen backbonelanguage modelingtrainable parametersfinetuningrouting modules
Authors
Valentijn Oldenburg, Floris de Kam, Bente Zuijdam, Lieve Eberson, Nicky van Zutphen, Stef de Wildt, Ivo Verhoeven
Abstract
Most parameter-efficient finetuning (PEFT) methods adapt weights or activations, thus leaving one of the key Transformer components unchanged: residual connections. This paper investigates Manifold-Constrained Hyper-Connections (mHC), a generalisation of residual connections, as a novel PEFT approach, wrapping frozen OLMo-2 backbones with learned residual routing modules. We find that mHC can finetune frozen Transformers, but that its role differs fundamentally from the original pre-training setting: in finetuning, fixing the residual mixing matrix to identity often improves performance. As a standalone PEFT method, mHC does not consistently outperform LoRA. However, at matched trainable parameter budgets, mHC+LoRA combinations improve language-modelling loss and show task-dependent benchmark gains at both 1B and 7B scale. Overall, our results identify residual routing as a distinct and promising novel PEFT axis.