Papers for

biotech companies

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Simple methods improve gene expression prediction from H&E images

Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?

Abstract: Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in several studies, but what they already solve and where additional complexity is needed remain unclear. We study this behavior through the structure of prediction error under the mean-squared error (MSE) objective. Differences in average expression across genes can account for a substantial part of aggregate prediction performance, while a key unresolved error lies in recovering variation within each slide. Decomposing MSE into slide-level and within-slide components, we find that the within-slide component has lower residual-normalized parameter sensitivity in controlled neural experiments. This motivates Component-Guided Loss (CGL), which increases supervision of the within-slide component. CGL-Linear is a closed-form affine instantiation that achieves overall state-of-the-art performance across HEST-1k cohorts and gene-panel sizes. The same within-slide supervision improves existing neural models. These results suggest that substantial gains can come from aligning the training objective with prediction-error structure rather than increasing model complexity.

Sat 26 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Predicting gene activity in tissue samples usually requires complex AI models, but the authors found that simpler methods can perform just as well by focusing on specific errors within each sample. They show that understanding the types of mistakes made during prediction leads to better training methods. This means that models can be improved not by being more complicated but by using better ways to learn from the data.
Open → 2609.32857v1

GyroNovo improves peptide sequencing by focusing on key missing fragments

GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for \textit{De Novo} Peptide Sequencing

Abstract: De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, and incomplete, leaving informative b- and y-ion fragments unobserved. Existing methods attempt to recover this missing evidence via latent-space imputation before autoregressive decoding. However, they typically treat imputation as a fixed reconstruction task, without considering which missing fragments are most relevant to decoder errors. Moreover, existing peak representations do not explicitly model mass differences between peaks, despite their fundamental importance. We introduce GyroNovo, a framework with two main contributions. First, we use decoder errors observed during training to adapt the imputation objective, prioritizing fragments associated with frequent decoding errors. We further use the decoder error distribution to construct easy and hard augmented views of each spectrum, enabling the decoder to learn under varying degrees of spectral corruption and missing-fragment severity. Second, we introduce a mass-aware inductive bias into self-attention by using rotary embeddings to encode pairwise mass differences between spectral peaks. Together, these components align missing-fragment recovery with decoder behavior while explicitly incorporating the mass relationships that underlie peptide fragmentation. At inference time, GyroNovo retains a standard encoder-imputer-decoder architecture and requires neither additional inputs nor auxiliary search procedures. Experiments on NovoBench show gains of about 9 percentage points in peptide-level precision and 7 percentage points in amino-acid-level precision over the state-of-the-art baseline. Code: https://github.com/UBC-NLP/gyronovo.

Thu 24 SeptMachine Learning
The gist
Peptide sequencing helps identify small protein pieces by analyzing their patterns in mass spectrometry data, but this data is often noisy and incomplete. The researchers created GyroNovo, a system that better guesses the missing pieces by focusing on parts most tied to errors during learning and by understanding the mass differences between fragments. Their approach helps the system learn even when data is messy and improves the accuracy of identifying peptides without needing a reference database. Testing showed GyroNovo performs significantly better than previous methods.
Open → 2609.30542v1