Papers for

biotech software developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Protein fitness prediction improves by combining models with evolutionary data

A General Harness for Protein Foundation Model Fitness Prediction

Abstract: Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their training corpora, while unreliable inputs can further distort fitness predictions. Family-specific evolutionary evidence and structural context can help address these limitations by providing complementary constraints on model scores, motivating VenusREM-Harness (VRH), a general, model-agnostic, training-free Retrieval-Enhanced Mutation harness. It fuses frozen model scores with multiple sequence alignment (MSA) evidence according to model uncertainty, then applies gated background correction and score shrinkage based on structural confidence and solvent exposure. Across 1,211 assays and 3.1 million measured variants from ProteinGym, VenusMutHub, and the newly curated viral benchmark VenusViroHub, all 71 configurations improve Spearman correlation on all 3 benchmarks by 0.073 on average, with broad gains across 5 metrics. Extended analyses relate retrieval gains to model-MSA preference differences, assess domain-level gains and immune-escape cases, and quantify computational speedups. Built with VRH, VenusREM2 is the first to rank highest in all function, taxon, MSA-depth, and mutation-depth categories, with a ProteinGym Average Spearman of 0.556, 0.038 above the prior best.

Mon 28 SeptArtificial Intelligence
The gist
Predicting how changes in proteins affect their function is important for designing new proteins. The paper describes a method that improves predictions by combining existing protein models with information about how proteins have evolved and their 3D structure. This approach helps correct biases and uncertainties in the models and leads to better accuracy across many tests. The authors built a tool called VenusREM2 that outperforms previous methods in predicting protein fitness.
Open → 2609.34654v1

Coding sequence optimization runs faster and uses less memory

SparseDesign: Scaling Exact Coding-Sequence Design

Abstract: Exact optimization of synonymous coding sequences under a joint folding-energy and codon-usage objective is limited by expensive dynamic-programming splits and large working sets. \textsc{SparseDesign} applies candidate sparsification to the multiloop recurrence of a Turner~2004 dangle-0 solver over a weighted codon automaton. A direct branch is retained only when it strictly improves on every partitionable or endpoint-unpaired realization of the same endpoint states. We prove equivalence to the dense recurrence in real arithmetic, under an explicit scalar branch-interface assumption. With $N$ automaton states, edge set $E$ and $Z$ retained candidates, multiloop work is $O(N^2+N|E|+NZ)$; worst-case time remains cubic for bounded-width automata and total memory remains quadratic. Endpoint ownership permits parallel candidate construction without locks. While synthetic stress families can benefit little from sparsification and exhibit near-quadratic candidate growth, natural proteins show substantial candidate-count reductions. In our 7,600-task campaign, the 2,000-protein human-table panel has median retention of only 3.53\% at $λ=0$ and 2.15\% at $λ=4$, corresponding to approximately 28.3-fold and 46.4-fold reductions relative to all feasible direct intervals. The primary performance experiments use an AMD EPYC 7313 server. For human Dp427c (11,031 nt, $λ=0$), 16-thread packed \textsc{SparseDesign} achieves five-run medians of 236.54 seconds wall-clock time and 14.43 GiB peak RSS. Compared with the single-thread local dense LinearDesign fork on the same server (4,912 seconds, 402.10 GiB RSS), this gives a 20.8-fold wall-clock speedup and a 27.9-fold peak-memory reduction. On a Core i9-14900KF commodity PC with 64 GiB RAM, the same input, layout and thread count achieve 126.42 seconds and 14.43 GiB RSS.

Sat 26 SeptData Structures and Algorithms
The gist
Designing DNA sequences that code for the same proteins but fold into specific shapes is very hard and uses lots of computer time and memory. The authors introduce SparseDesign, a new way to cut down the number of possibilities the program checks without missing better solutions. This method speeds up the design process a lot and cuts memory use by more than 20 times on real protein examples. SparseDesign works well in practice, especially for natural proteins, enabling faster and more efficient genetic sequence design.
Open → 2609.32308v1

Graph neural network predicts protein membrane structure from atoms

Predicting Transmembrane Protein Topology from 3D Structure

Abstract: This paper presents a novel approach to infer protein topology using the state-of-the-art graph neural network (GNN), SchNet. The model is trained on the same dataset used to develop the recent DeepTMHMM model with 5-fold cross-validation. Unlike the conventional approaches based on using only the protein sequences or the $α$-carbons as features, we have decoded our classifier in this way, so all atom-level embeddings are used. Without applying any pre-trained weight, the final results have shown great potential that GNNs can be used for topological predictions.

Thu 24 SeptArtificial Intelligence
The gist
Figuring out how proteins fit into cell membranes is important for biology and medicine. This paper shows a new way to predict these protein shapes using a graph-based AI model called SchNet, which looks at all the atoms in the protein. The authors trained their model on data previously used for another method called DeepTMHMM and got promising results. Their approach is different because it uses the full 3D structure instead of just protein sequences or a few atoms. This suggests that graph neural networks can be helpful in understanding protein membrane topologies.
Open → 2609.30446v1

MorphoOrgaAgent automates organoid analysis with natural language input

MorphoOrgaAgent: A Foundation-Model-Based Multi-Agent System for Autonomous Organoid Analysis

Abstract: Organoids are three-dimensional tissue models whose morphology provides important insights into tumor development, disease progression, and drug testing. Extracting these morphological features relies heavily on manual segmentation, which is time-consuming and labor-intensive. Furthermore, performing quantitative statistical analysis typically requires custom coding skills and a mathematical background, presenting a major barrier for experimental biologists. To address these challenges, we introduce MorphoOrgaAgent, a multi-agent framework that achieves zero-shot organoid segmentation, automated data analysis, and report generation based on natural language input. The framework consists mainly of three core components: a TaskUnderstandingAgent that identifies requested measurements and visualization types; a hybrid segmentation module that combines Cellpose-derived geometric prompts with text prompts to guide SAM3 for zero-shot organoid instance segmentation; and a ReportAgent that computes quantitative metrics and compiles them alongside generated visualizations into a structured report. We further introduce MorphoOrgaVQA, a benchmark designed for quantitative evaluation of agent systems in organoid morphology analysis. Experimental results demonstrate that MorphoOrgaAgent handles both explicit and descriptive user requests, produces measurements closely matching ground truth, and generates complete analysis reports without requiring manual programming. The complete source code and MorphoOrgaVQA benchmark are publicly available at https://github.com/peng-lab/MorphoOrgaAgent.

Tue 8 SeptMultiagent SystemsComputer Vision and Pattern Recognition
The gist
Studying tiny 3D cell clusters called organoids is important for understanding diseases and testing drugs, but analyzing their shapes usually takes a lot of manual work and coding skills. The authors created MorphoOrgaAgent, a computer system that understands what measurements and images a user wants from simple language, then automatically segments organoids from microscope pictures and generates detailed reports. This system works without needing prior training (zero-shot) and matches expert measurements closely. The authors also provide a benchmark to test such systems, and they share their code for others to use.
Open → 2609.08696v1