Spinet predicts protein sequences accounting for motion dynamics

SPINET: Sheaf Protein Inverse Folding Network

Machine Learning

Summary

Proteins change shape while working, but most methods guess their sequences from a single shape. The authors created SPINET, a new tool that looks at how proteins move over time to better predict their building blocks. It uses a math-based approach to understand interactions in each snapshot and combines these over time to make predictions. SPINET performed better than previous methods in matching known protein sequences and shapes.

What this means in practice

  • For protein engineers: Design proteins that change shape as needed by predicting sequences compatible with dynamic structures along motion trajectories.
  • For biotechnology developers: Create drug candidates or enzymes with specific dynamic behaviors by leveraging sequence predictions conditioned on protein motion data.

Authors

Jens Lundsgaard, Colin Mikulski, Zhixuan Yan, Dhananjay Bhaskar

Abstract

Proteins change shape as they function, yet most inverse folding models predict amino acid sequences from a single, fixed backbone. A central challenge in protein engineering is to design proteins that undergo specific motions, which requires accounting for how their structures change over time. This motivates inverse protein folding conditioned on protein motion. We introduce SPINET, which predicts sequences from molecular dynamics trajectories. It uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass. We evaluate SPINET on mdCATH and ATLAS, where it outperforms all evaluated static and ensemble baselines in sequence recovery. On mdCATH, it achieves 56.7% top-1 recovery, compared with 44.5% for the strongest static baseline and 40.7% for the strongest ensemble baseline. We also evaluate whether the predicted sequences are compatible with conformations sampled along the target trajectory. On mdCATH, they achieve a median TM-score of 0.760, and structural recovery favors target conformations over unrelated decoys for 99.5% of test domains.