Papers for

protein engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Few-step generative models aligned using direct preference optimization

FestDPO: Few-step Generator Alignment with Direct Preference Optimization

Abstract: Few-step generative models can generate high-fidelity samples within a few function evaluations. Despite this efficiency, generated samples may not exhibit desirable properties. When these properties are difficult to encode as an explicit reward function, direct preference optimization (DPO) can align generative models using pairwise preference feedback without training a separate reward model. However, extending DPO to few-step generative models is challenging because few-step generative models are generally implicit, making the likelihood evaluation required by DPO intractable. To address this challenge, we introduce Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples. By exploiting the fast sampling capabilities of few-step generative models, our approach makes sample-based approximation of DPO loss computationally feasible. Furthermore, the sample-based formulation makes FestDPO agnostic to the model family and sampling procedure. Our toy experiment demonstrates that FestDPO matches the reward-tilted target distribution across four few-step generators. For real-world tasks, we evaluate FestDPO in two domains: text-to-image generation and protein backbone generation. In text-to-image generation, FestDPO outperforms preference optimization baselines in both win rates against the base models and human evaluation scores. In protein backbone generation, it achieves a higher $β$-sheet fraction and better structural designability than the baselines.

Mon 28 SeptMachine Learning
The gist
Some computer programs can quickly create things like images or protein shapes, but these creations might not always meet what we want. The authors worked on a way to improve these programs by teaching them with human preferences directly, without needing to guess scores in between. They developed a new method called FestDPO that can handle these quick programs even though they don’t easily show how likely their outputs are. Tests showed FestDPO can better align creations with human preferences in pictures and proteins.
Open → 2609.34673v1

Protein fitness prediction improves by combining models with evolutionary data

A General Harness for Protein Foundation Model Fitness Prediction

Abstract: Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their training corpora, while unreliable inputs can further distort fitness predictions. Family-specific evolutionary evidence and structural context can help address these limitations by providing complementary constraints on model scores, motivating VenusREM-Harness (VRH), a general, model-agnostic, training-free Retrieval-Enhanced Mutation harness. It fuses frozen model scores with multiple sequence alignment (MSA) evidence according to model uncertainty, then applies gated background correction and score shrinkage based on structural confidence and solvent exposure. Across 1,211 assays and 3.1 million measured variants from ProteinGym, VenusMutHub, and the newly curated viral benchmark VenusViroHub, all 71 configurations improve Spearman correlation on all 3 benchmarks by 0.073 on average, with broad gains across 5 metrics. Extended analyses relate retrieval gains to model-MSA preference differences, assess domain-level gains and immune-escape cases, and quantify computational speedups. Built with VRH, VenusREM2 is the first to rank highest in all function, taxon, MSA-depth, and mutation-depth categories, with a ProteinGym Average Spearman of 0.556, 0.038 above the prior best.

Mon 28 SeptArtificial Intelligence
The gist
Predicting how changes in proteins affect their function is important for designing new proteins. The paper describes a method that improves predictions by combining existing protein models with information about how proteins have evolved and their 3D structure. This approach helps correct biases and uncertainties in the models and leads to better accuracy across many tests. The authors built a tool called VenusREM2 that outperforms previous methods in predicting protein fitness.
Open → 2609.34654v1

Spinet predicts protein sequences accounting for motion dynamics

SPINET: Sheaf Protein Inverse Folding Network

Abstract: Proteins change shape as they function, yet most inverse folding models predict amino acid sequences from a single, fixed backbone. A central challenge in protein engineering is to design proteins that undergo specific motions, which requires accounting for how their structures change over time. This motivates inverse protein folding conditioned on protein motion. We introduce SPINET, which predicts sequences from molecular dynamics trajectories. It uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass. We evaluate SPINET on mdCATH and ATLAS, where it outperforms all evaluated static and ensemble baselines in sequence recovery. On mdCATH, it achieves 56.7% top-1 recovery, compared with 44.5% for the strongest static baseline and 40.7% for the strongest ensemble baseline. We also evaluate whether the predicted sequences are compatible with conformations sampled along the target trajectory. On mdCATH, they achieve a median TM-score of 0.760, and structural recovery favors target conformations over unrelated decoys for 99.5% of test domains.

Mon 28 SeptMachine Learning
The gist
Proteins change shape while working, but most methods guess their sequences from a single shape. The authors created SPINET, a new tool that looks at how proteins move over time to better predict their building blocks. It uses a math-based approach to understand interactions in each snapshot and combines these over time to make predictions. SPINET performed better than previous methods in matching known protein sequences and shapes.
Open → 2609.34153v1

Protein design improves stability across shape changes

Robust Biomolecular Complex Design Across Protein Conformational Landscapes

Abstract: Proteins populate conformational ensembles, yet structure-based biomolecular design typically optimizes candidates against a single target conformation. Consequently, a candidate that fits one state can lose favorable interactions or develop steric clashes when the target adopts another. We introduce FlexEvo, a model-agnostic evolutionary framework that adapts candidates once at inference time from a single target conformation to improve compatibility with alternative natural conformations unseen during adaptation, without retraining the source model or requiring a conformational ensemble. FlexEvo casts cross-state adaptation as geometry-constrained bi-objective optimization, balancing preservation of input-state interactions against robustness to plausible conformational perturbations. To limit the search space and reduce invalid structural edits, geometry-derived FlexBoxes define protected anchor regions, adaptable regions for local exploration, and forbidden regions for clash avoidance. A unified all-atom representation supports topology-preserving adaptation across diverse binder categories, while Pareto selection preserves nondominated candidates across the two objectives. We evaluate FlexEvo across multiple generation baselines and nine representative binder categories spanning diverse molecular sizes and structural topologies. FlexEvo reduces the category-balanced mean relative performance degradation from 47.8% to 4.4%, while adding only 1.4--3.1 minutes of adaptation per sample. These results establish single-state inference-time adaptation as a practical route toward robust biomolecular complex design across protein conformational landscapes.

Sun 27 SeptArtificial Intelligence
The gist
Proteins can change their shapes, but designing molecules to fit them usually focuses on just one shape. This causes problems when the protein shifts and the designed molecule no longer fits well. The authors created FlexEvo, a way to tweak protein designs quickly so they work well across different shapes without needing to redesign from scratch. It balances keeping the original good features and avoiding clashes when the protein moves. Testing shows this method greatly reduces loss of performance while adding only a little bit of extra design time.
Open → 2609.33726v1