Simple methods improve gene expression prediction from H&E images
Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?
Computer Vision and Pattern RecognitionMachine Learning
Summary
Predicting gene activity in tissue samples usually requires complex AI models, but the authors found that simpler methods can perform just as well by focusing on specific errors within each sample. They show that understanding the types of mistakes made during prediction leads to better training methods. This means that models can be improved not by being more complicated but by using better ways to learn from the data.
What this means in practice
- •For medical image analysts: Predict spatial gene expression from routine H&E images with improved accuracy using Component-Guided Loss methods.
- •For biotech companies: Develop scalable spatial molecular profiling tools with simpler models that reduce complexity while maintaining prediction quality.$Commercial implications: Enables creation of cost-effective spatial transcriptomics platforms for research and diagnostics by improving model training strategies.
Authors
Duc T. Nguyen, Thanh Ha Do, Phuong M. Cao, Hieu Pham
Abstract
Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in several studies, but what they already solve and where additional complexity is needed remain unclear. We study this behavior through the structure of prediction error under the mean-squared error (MSE) objective. Differences in average expression across genes can account for a substantial part of aggregate prediction performance, while a key unresolved error lies in recovering variation within each slide. Decomposing MSE into slide-level and within-slide components, we find that the within-slide component has lower residual-normalized parameter sensitivity in controlled neural experiments. This motivates Component-Guided Loss (CGL), which increases supervision of the within-slide component. CGL-Linear is a closed-form affine instantiation that achieves overall state-of-the-art performance across HEST-1k cohorts and gene-panel sizes. The same within-slide supervision improves existing neural models. These results suggest that substantial gains can come from aligning the training objective with prediction-error structure rather than increasing model complexity.