Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling
2026-08-17 • Artificial Intelligence
Artificial IntelligenceMachine Learning
AI summaryⓘ
The authors explain how predicting protein structures has improved a lot thanks to deep learning, which helps us understand how proteins work. They review how methods have evolved from using evolutionary information and simple features to more advanced models like AlphaFold2 and RoseTTAFold, which learn from large data directly. The review also shows how the field moved from predicting single proteins to modeling protein complexes and designing new proteins with generative models. This helps clarify how changes in methods have affected what these tools can do and their practical uses.
protein structure predictiondeep learningmultiple sequence alignment (MSA)AlphaFold2RoseTTAFoldprotein complexesgenerative modelingstructural biologyevolutionary couplingmolecular function
Authors
Wengan He, Yongsheng Luo, Lihong Jiang, Wenhui Xu, Yu Li
Abstract
Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provides a basis for mechanistic interpretation. Recent advances in deep learning have transformed the field from multiple sequence alignment (MSA)-driven monomer folding into broader frameworks capable of modeling protein complexes and increasingly heterogeneous molecular systems. Existing reviews have summarized this progress from the perspectives of representative models, application domains, and protein design. Building on these efforts, this review focuses on the methodological evolution of the field itself. It examines recent developments through three closely related dimensions: representations and data, architectures and learning strategies, and confidence and evaluation. Within this perspective, the field is organized into four methodological phases and three cross-cutting transitions: from explicit evolutionary coupling features and early contact prediction to learned sequence representations in AlphaFold2, RoseTTAFold, and ESMFold; from protein-only monomer folding to increasingly integrated modeling of heterogeneous molecular systems in AlphaFold-Multimer, RoseTTAFoldNA, and AlphaFold3; and, more recently, from prediction-oriented structure inference to design-oriented generative modeling in RFdiffusion and related frameworks. This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.