Protein language model method improves acid-loving protein identification

MI-PEFT: Mixture-of-Experts Integrated Parameter-Efficient Fine-Tuning Protein Language Models Improves Acidophilic Proteins Classification

Machine Learning

Summary

Identifying proteins that work well in very acidic environments is important for many industries but usually takes a lot of time and experiments. The researchers developed a new computer method called MI-PEFT that uses advanced protein language models to make this identification faster and more efficient. Their method combines clever ways to fine-tune the models while handling difficult data problems like when some protein types are much rarer than others. Tests showed that MI-PEFT can better recognize acid-loving proteins and keep important information from large pre-trained models. This makes the process more accurate and less time-consuming.

protein language modelsacidophilic proteinsparameter-efficient fine-tuningmixture-of-expertsESM C-600MLoRAclass imbalanceDeepSeekMoEprotein classificationpretrained representations

Authors

Honghan Shen

Abstract

Acidophilic proteins that remain stable and functional under highly acidic conditions, are important for industrial biocatalysis, acid-related bioprocessing, and the discovery of acid-stable enzymes. However, their identification relies heavily on time-consuming experimental screening methods. With the rapid growth of protein sequence databases, the need for computational identification methods that are both accurate and efficient has become stronger. The emergence of protein language models (PLMs) has significantly improved the sequence representation of downstream biological prediction tasks. This paper proposes MI-PEFT, a mixture-of-experts integrated parameter-efficient fine-tuning framework. Built on the ESM C-600M backbone, the framework incorporates LoRA-based PEFT methods and a DeepSeekMoE-based classification head to resolve the limitations of PEFT and significantly improve computational efficiency. Notably, this task is characterized by a significant class imbalance in the dataset, making high specificity particularly challenging. The experimental results demonstrate that MI-PEFT on PLMs, especially {\text{C}}^{\text{3}}\text{A}, serves as an efficient tool for identifying acidophilic proteins and a constrained pathway that helps resolve class-imbalance by preserving the pretrained representations.