Idiom model and rl-sae method generate and control disordered protein sequences

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Machine Learning

Summary

Intrinsically disordered protein regions (IDRs) are flexible parts of proteins important for many cell functions but are hard to design because they don't have fixed structures. The authors created IDiom, a language model trained specifically on these disordered regions, to generate realistic sequences. They also developed a method called RL-SAE that helps control specific sequence features linked to function by rewarding the model when it generates desired patterns. This approach produces sequences that better match desired biological activities than previous methods. Their tools allow more interpretable and targeted design of these flexible protein regions.

What this means in practice

  • For protein engineers: Generate and customize disordered protein sequences with controlled biological features to design proteins with targeted cellular functions.
  • For biotechnology developers: Incorporate RL-SAE to improve protein design platforms by enabling interpretable and programmable control over disorder-related sequence features.

Authors

Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff

Abstract

Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at https://github.com/rotskoff-group/idiom.