Pathology role-guided experts improve whole slide image classification
Role-Guided MOE for Encoder-Level Pathology Representation Learning in WSI Classification
Computer Vision and Pattern Recognition
Summary
Classifying whole slide images (WSIs) of tissue is important for diagnosing diseases. The authors found that current models often use fixed feature extractors that don’t adapt well to unique tissue patterns. They created a mixture-of-experts method that guides parts of the model to specialize in different pathological roles, improving how tissue patterns are represented. This method helps the model work better, especially with limited data, and showed consistent improvement on multiple pathology datasets.
What this means in practice
- •For medical imaging developers: Improve tissue feature extraction in pathology slide analysis tools to boost diagnostic accuracy using role-guided mixture-of-experts models.
- •For clinical ai teams: Enhance whole slide image classifiers to better adapt to diverse tissue types in limited data scenarios for more reliable cancer diagnosis support.
Authors
Xinyu Ma, Xing Yang, Hongtao Jin, Guoquan Zhang, Shijie Zhang, Yu Zhang, Xitong Li
Abstract
Whole slide image classification is a fundamental task in computational pathology, where patch representation quality directly affects downstream aggregation and slide-level discriminability. Pathology foundation models are widely adopted as frozen feature extractors for WSI classification; however, their fixed encoders may produce representations insufficiently adapted to target-specific tissue patterns and discriminative cues. Fine-tuning can improve target adaptation, but introduces a trade-off between pathology-specific representation capacity and adaptation efficiency, particularly in data-scarce settings. To address this, we propose a pathology role-guided mixture-of-experts feed-forward network (MoE-FFN) framework for efficient encoder-level representation learning. We design a two-stage training paradigm to establish and adapt pathology-aware expert specialization. In source-domain expert initialization, pathology-specific priors are distilled from a frozen Virchow2 teacher into a lightweight DINOv2-small student, while role prototypes serve as weak pathological anchors to encourage distinct expert functions. MoE-FFN blocks are introduced into selected high-level transformer layers to provide transformation diversity for heterogeneous pathological patterns. In target-domain adaptation, the initialized experts are refined through asymmetric prototype-guided optimization, enhancing task-relevant positive evidence and separating confusable hard negatives. The resulting encoder extracts offline patch representations that can be directly integrated with standard MIL aggregators. Experiments on the public BRACS dataset and a private PAROTID WSI dataset across five representative backbones demonstrate consistent improvements over the strongest baseline.