On-device language models personalize efficiently with loRA-generating hypernetworks

LoRA-generating hypernetworks for efficient on-device LLM generative personalization

Machine Learning

Summary

Mobile phones can't run huge language models easily, so improving them to better fit each user is hard. The authors introduce a new way to customize these models right on your device using a small extra network that creates a personalized adjustment called LoRA. This method is faster and uses less memory than other ways and changes the model itself instead of just feeding it more information. They tested their approach on difficult text tasks and found it works better than previous methods.

What this means in practice

  • For mobile app developers: Create personalized mobile language apps that improve text generation quality without heavy computing or storage demands.$Commercial implications: Enables practical deployment of personalized LLM features on phones, unlocking new app products with enhanced user-tailored text generation.
  • For embedded systems engineers: Build low-latency on-device language models that adapt to user behavior efficiently, saving compute costs in resource-limited environments.

Authors

Sean Augenstein, Li Ding, Jihwan Lee, Keith Rush, Andrey Zhmoginov

Abstract

On-device large language models (`LLMs'), e.g. running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains highly impactful. At the same time, their personal nature (i.e., the close coupling to a particular user) means that a given on-device LLM tends to be used in similar, predictable patterns over the course of time. This paper presents a novel method for personalizing on-device LLMs. It trains a hypernetwork to map a user's context tokens to a low-rank adaptation (`LoRA') well-suited to that user. Once the trained common artifacts are deployed to users' devices, each user uses the hypernetwork to synthesize (entirely on device) a personalized LoRA. This approach blends the benefits while avoiding the drawbacks of two existing approaches to LLM customization: in-context learning (`ICL') and parameter-efficient fine-tuning (`PEFT'). Like ICL (and unlike PEFT), the on-device phase of our approach is computationally feasible, requiring only forward passes through neural networks. Like PEFT (and unlike ICL), our approach modifies the `target' base LLM via weights (the LoRA), avoiding negative consequences (e.g. increased latency) associated with extending the input sequence. Our approach is particularly well-suited to the mobile device regime. Apart from the on-device compute and latency benefits mentioned, it also requires minimal additional storage, as internally its architecture partly leverages the same LLM weights as belong to the target LLM to be personalized. We demonstrate the benefits of LoRA-generating hypernetworks on several representative personalization datasets, comparing against baselines like ICL and PEFT. Of note, our personalization experiments focus on more challenging and less studied long-form text generation tasks.