Untied embeddings improve private training of large language models
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
Machine Learning
Summary
Training large language models with privacy protections is important but challenging. The authors studied how a common design choice called weight tying affects private training methods. They found that not sharing these weights led to better accuracy and much lower memory usage under privacy-preserving training. This suggests that standard model designs may need to change when privacy is required. Their work helps make private model training more efficient and effective.
What this means in practice
- •For machine learning engineers: Improve memory usage and accuracy when privately fine-tuning GPT-style models by untying input and output embeddings.
- •For cloud service providers: Offer more scalable privacy-preserving model fine-tuning services by adopting embedding designs compatible with memory-efficient DP training methods.
Authors
Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine
Abstract
Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.