Protein fitness prediction improves by combining models with evolutionary data
A General Harness for Protein Foundation Model Fitness Prediction
Artificial Intelligence
Summary
Predicting how changes in proteins affect their function is important for designing new proteins. The paper describes a method that improves predictions by combining existing protein models with information about how proteins have evolved and their 3D structure. This approach helps correct biases and uncertainties in the models and leads to better accuracy across many tests. The authors built a tool called VenusREM2 that outperforms previous methods in predicting protein fitness.
What this means in practice
- •For protein engineers: Improve the accuracy of predicting how mutations affect protein function by integrating evolutionary and structural data with existing models.
- •For biotech software developers: Implement a model-agnostic tool that enhances protein mutation scoring by combining frozen model outputs with sequence alignments and structural corrections.
Authors
Yang Tan, Qijia Tian, Gangyu Sun, Bozitao Zhong, Mingchen Li, Yuanxi Yu, Nanqing Dong, Liang Hong
Abstract
Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their training corpora, while unreliable inputs can further distort fitness predictions. Family-specific evolutionary evidence and structural context can help address these limitations by providing complementary constraints on model scores, motivating VenusREM-Harness (VRH), a general, model-agnostic, training-free Retrieval-Enhanced Mutation harness. It fuses frozen model scores with multiple sequence alignment (MSA) evidence according to model uncertainty, then applies gated background correction and score shrinkage based on structural confidence and solvent exposure. Across 1,211 assays and 3.1 million measured variants from ProteinGym, VenusMutHub, and the newly curated viral benchmark VenusViroHub, all 71 configurations improve Spearman correlation on all 3 benchmarks by 0.073 on average, with broad gains across 5 metrics. Extended analyses relate retrieval gains to model-MSA preference differences, assess domain-level gains and immune-escape cases, and quantify computational speedups. Built with VRH, VenusREM2 is the first to rank highest in all function, taxon, MSA-depth, and mutation-depth categories, with a ProteinGym Average Spearman of 0.556, 0.038 above the prior best.