Papers for
medical image analysts
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Preference guidance improves open-vocabulary image segmentation in new domains
Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement
Abstract: Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial inspection, where dense pixel-level masks for adaptation are costly to obtain and require domain-specific expertise. We propose a preference-guided adaptation framework that replaces dense mask supervision with binary preferences. We observe that different prompt templates produce systematically different segmentations for the same image, a phenomenon we call prompt disagreement, and we repurpose it as a built-in source of preference supervision. Building on this, we mine localized preference queries from regions of high cross-template uncertainty, and adapt the OVSS model with Region-Localized Preference Optimization (RLPO) together with consistency regularization that stabilizes updates outside the queried region. Across extensive experiments on the MESS benchmark, the proposed method achieves consistent gains across diverse OVSS backbones without any pixel-level annotation, and remains effective under noisy preferences. Our code is available at https://github.com/blue-531/pref-ovss.
Model uncertainty does not match human ambiguity in vision tasks
Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks
Abstract: Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement ($ρ= 0.24--0.55$), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.
Whole-slide image analysis models tumor microenvironment interactions dynamically
Modeling Whole-Slide Images as Dynamic Tumor Microenvironment Fields
Abstract: Due to the gigapixel-scale nature of whole-slide images (WSIs), weakly supervised WSI analysis is commonly formulated as a multiple instance learning (MIL) problem, where patch-level features are aggregated into slide-level representations. However, diagnostic and prognostic evidence often arises from spatially coherent tumor microenvironment regions and their interactions, rather than isolated patches alone. Existing patch-level or static region-based methods usually overlook how tissue regions should be adaptively formed and subsequently evolved through microenvironment interactions across heterogeneous boundaries. In this paper, we propose Concept-Guided Tumor Microenvironment Evolution (TMEvolve), a reaction-diffusion-inspired framework that models WSIs as latent tumor microenvironment fields over discrete patch graphs. TMEvolve instantiates this view as a learnable graph-discretized evolution process over patch neighborhoods. It first forms adaptive soft tissue regions as coherent microenvironment units, then performs pseudo-time evolution through two complementary local dynamics: intra-region diffusion, which stabilizes latent states within coherent tissue compartments, and concept-guided boundary flux, which propagates visual feature signals and language-derived concept signals across heterogeneous region interfaces. The evolved microenvironment regions are finally aggregated for slide-level prediction. We evaluate TMEvolve on six datasets across three weakly supervised WSI tasks: survival prediction, gene expression prediction, and histological subtype classification. TMEvolve consistently improves over representative MIL methods, pathology foundation models, and concept-guided baselines. Ablation studies and visualizations further support the effectiveness and interpretability of TMEvolve, highlighting the value of dynamic region modeling and boundary interaction.
Quantum method improves medical image retrieval with fewer parameters
QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG
Abstract: Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent retrieval adaptation (QuPID) repairs this by making the circuit input-dependent through data re-uploading and by comparing measurement readouts, vectors of local Pauli expectations, rather than states. The result is a small readout for adapting frozen image features to a local archive with limited data: training simulates the circuit classically, and inference runs on a GPU with fixed learned parameters. We characterize the class as a structured factorization of input-modulated quadratic feature maps, bound the frequency support of its re-uploading channel, and give a parameter-count generalization bound that motivates its small budget. Under a shared frozen backbone and a label-free protocol, QuPID's 60 parameters give higher precision-at-5 (P@5) on ChestX-ray14 and MURA than frozen medical encoders, and than adapters and low-rank adaptation (LoRA) with up to 5.25 million trainable parameters. On ChestX-ray14, the P@5 gain over the frozen encoder is +0.116, the lead over retuned adapters is widest at 512 adaptation examples (+0.040), and the full-budget margin over an equally compact classical rotation-plane head is +0.023 with a 95% interval excluding zero. Medical imaging is the primary testbed; the pattern recurs on two non-medical benchmarks, in report generation, and under simulated gate noise and finite-shot readout.
Efficient image restoration with flexible learned degradation operators
Perturb-and-Solve: Efficient Learned-Operator Conditioning for Latent Diffusion Inverse Problems
Abstract: Latent diffusion models serve as powerful priors for solving inverse problems in image restoration, such as deblurring, inpainting, and super-resolution. Current methods have a trade-off between generality and efficiency. Solvers that are restricted to a fixed set of degradation operators are fast and efficient. Methods that support arbitrary degradation operators are slow and require gradients through the diffusion network. To break this bottleneck, we introduce PASEO (Perturb-And-Solve for Efficient Operator conditioning), a method that uses a small (1M parameters) learned network to degrade diffusion model predictions in latent space. PASEO supports learned degradation operators without back-propagating through the diffusion network. We efficiently sample reconstructions from an approximate posterior by combining the diffusion model's prediction with the observed image. We do this by adding noise and solving linear equations based on a local linear approximation of the learned network, without building or inverting large covariance matrices. Across super-resolution, deblurring, and inpainting on FFHQ and COCO, PASEO achieves strong perceptual quality while running up to 9x faster and using up to 34% less peak memory than the tested baselines, with the same or fewer model evaluations.
Simple methods improve gene expression prediction from H&E images
Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?
Abstract: Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in several studies, but what they already solve and where additional complexity is needed remain unclear. We study this behavior through the structure of prediction error under the mean-squared error (MSE) objective. Differences in average expression across genes can account for a substantial part of aggregate prediction performance, while a key unresolved error lies in recovering variation within each slide. Decomposing MSE into slide-level and within-slide components, we find that the within-slide component has lower residual-normalized parameter sensitivity in controlled neural experiments. This motivates Component-Guided Loss (CGL), which increases supervision of the within-slide component. CGL-Linear is a closed-form affine instantiation that achieves overall state-of-the-art performance across HEST-1k cohorts and gene-panel sizes. The same within-slide supervision improves existing neural models. These results suggest that substantial gains can come from aligning the training objective with prediction-error structure rather than increasing model complexity.
Ultrasound image clarity improved by new self-supervised despeckling method
Mask2Restore: Self-Supervised Ultrasound Despeckling via Inpainting
Abstract: Medical ultrasound (US) is inherently degraded by speckle, a granular interference pattern that is often treated as a complex form of noise in image restoration. However, unlike random noise, US speckle originates from coherent scattering within tissue and is therefore highly spatially dependent and deterministic under fixed acquisition conditions, making US speckle suppression fundamentally different from natural image denoising. Because speckle-free US targets are unavailable in practice, self-supervised denoising is necessary. Blind-spot networks (BSN) are the dominant self-supervised paradigm for natural images, but their pixel-wise masking strategy assumes spatially independent noise, an assumption poorly matched to US speckle, which is spatially correlated over multiple pixels rather than pixel-wise independent. To address this mismatch, we propose Mask2Restore, a self-supervised US despeckling framework that reformulates despeckling as contextual inpainting with block-wise masking on single noisy images. Unlike pixel-wise BSN masking, block-wise masking addresses this multi-pixel speckle correlation by removing locally correlated speckle neighborhoods and shifting the reconstruction cues used by the network from adjacent speckle correlations to broader anatomical context. We further introduce cross-resolution context regularization (CRCR), which suppresses residual speckle bias by enforcing consistency across multi-resolution predictions. Experiments on simulated and in vivo carotid US, unseen fine-structure cases, and downstream cardiac segmentation demonstrate improved speckle-detail trade-offs, better preservation of fine anatomical structures, and practical value for subsequent image analysis.
Data scale factors affect brain image model training outcomes
Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture
Abstract: Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain histology, where a source is an individual brain, and a sample is an image patch at a specific spatial location. Across 93 controlled pretraining runs of a contrastive model that uses spatial proximity for supervision, we vary data allocation, compute, and model capacity over 11.6 million spatially anchored image patches from 21 human brains. Performance improves with more unique samples, broader spatial coverage, additional compute, and larger model capacity. At fixed sample count, distributing samples across one to 18 subjects produces no detectable improvement, even though representations generalize substantially better to subjects encountered during pretraining. Inter-subject variation therefore strongly affects generalization, but additional subjects provide no benefit when a fixed sample budget is distributed across more sources. These results establish sample count, source diversity, and spatial coverage as distinct axes of data scaling in spatially structured representation learning.
Image recovery improves using neural priors with Fourier phase data
Image Reconstruction from Phase with Untrained Neural Priors
Abstract: Fourier phase encodes important spatial image structure, but recovering an image without measured spectral magnitude requires additional constraints and leaves absolute intensity ambiguous. We propose a projection-based two-stage framework that combines Fourier-phase and spatial-support constraints with an image-specific neural prior. The first stage alternates constraint enforcement with regularized neural-prior updates, while the second performs phase/support refinement alone with guaranteed convergence. We evaluate two neural-prior implementations on the same 77 microscopy images and compare them with a constraint-only baseline. After 500 final refinement passes, the best-performing variant achieves 31.41 dB pooled PSNR, 35.75 dB mean PSNR, and 0.9531 mean SSIM, improving pooled PSNR by 1.51~dB and reducing pooled MSE by 29.3% relative to the baseline. The results demonstrate the benefit of combining neural guidance with explicit constraint refinement at the evaluated iteration budget, while showing that lower phase residual alone does not guarantee greater reconstruction accuracy.
Physics informed method reduces shadows in fetal ultrasound imaging
Shadow Reduction in Ultrasound Imaging Using Differentiable Simulation and Radiance Field Decomposition
Abstract: Acoustic shadows from bone and other highly attenuating tissues obscure clinically important structures in ultrasound. In fetal brain imaging, skull-induced artefacts disproportionately degrade the hemisphere closer to the transducer (proximal), limiting symmetric assessment of the two hemispheres. Existing correction methods require raw scanner data, impose restrictive assumptions on tissue properties, or rely on generative models that may hallucinate anatomy. We present RFlash, a physics-informed post-processing method that decomposes beamformed ultrasound images into explicit attenuation and scatter-intensity maps using a differentiable radiance-field formulation of image formation. Attenuation-adaptive re-rendering then removes the dependence of the signal at each depth on the intervening tissue, equivalent to virtually advancing the transducer into the tissue. Across 1,261 3D fetal brain volumes, 143 real 2D curvilinear abdominal scans, and 1,200 simulated 2D linear-probe liver scans, RFlash reduces shadow-related intensity differences more effectively than classical Hughes-Duck attenuation correction. For a gestational-age model trained on the distal hemisphere (further from the transducer) and applied to the proximal hemisphere, prediction error decreases by 5.1 days (40%) relative to the original images. The estimated attenuation maps also yield shadow-confidence maps that improve random-forest bone-shadow segmentation over the image alone and receive greater SHAP importance than an existing neural confidence-map baseline, suggesting greater physical consistency. RFlash requires neither hardware modification nor access to raw scanner data and supports 2D and 3D acquisitions with linear and curvilinear probes, making it widely applicable allowing clinicians to use our method on their already acquired scanners and images.
Polyp image synthesis improves colonoscopy data with adaptive mucosal context
Beyond the Foreground: FOV-Aware Polyp Image Synthesis via Lesion-Guided Adaptive Mucosal Context Propagation
Abstract: Synthetic image and mask pairs can alleviate scarce colonoscopy annotations, but realistic synthesis requires preserving the supplied lesion while generating compatible mucosa. Existing foreground-guided methods treat all non-foreground pixels as background and rely mainly on local integration. Directly applying them to colonoscopy causes two problems: non-mucosal black regions contaminate generated tissue, and local reasoning produces inconsistent mucosal texture and illumination. We propose LAMP, the first foreground-guided framework for polyp image synthesis based on lesion-guided adaptive mucosal context propagation. LAMP explicitly separates the lesion, valid mucosa, and camera exterior using a field-of-view (FOV) mask. Lesion-to-Mucosa cross-attention extracts lesion appearance conditions for valid-mucosa locations, while FOV-constrained multidirectional Vision Receptance Weighted Key Value propagates them over legal tissue support. An adaptive gate then controls their residual fusion into the diffusion U-Net. Extensive experiments on five polyp datasets demonstrate that LAMP substantially outperforms existing methods in overall generation quality and consistently improves five downstream segmentation models. Our code will be released at https://github.com/wangtong627/LAMP.
Normal kidney images help detect rare glomerulus abnormalities accurately
Seeing Abnormal from Normal: Glomerular Abnormality in Representations of Normal Renal Morphology
Abstract: Fine-grained evaluation of glomerular pathology must distinguish normal glomeruli from abnormalities such as global and segmental glomerulosclerosis, obsolescent, ischemic, solidified, disappearing, and atubular glomeruli. Supervised classification requires labeled examples of every category, which is impractical when subtypes are rare or absent from the training cohort. One-class anomaly detection offers an alternative by modeling normal data and scoring deviations, allowing previously unseen abnormalities to be detected. We use the frozen residual U-Net backbone of Omni-Seg, pretrained to segment structurally normal renal primitives without abnormal-subtype labels. We propose NoRDeC (Normal-Reference Detection and Characterization), a framework combining Mahalanobis normal-reference scoring with layer-wise representation analysis to determine whether and where glomerular pathology is encoded, how spatial aggregation affects detection, and whether abnormalities alter inter-layer relationships differently. Using glomerular images from two institutions, we evaluate backbone layers and aggregation strategies, compare NoRDeC with PaDiM and PatchCore, and analyze representations using centered kernel alignment (CKA). Layer 4 with Center-70 aggregation achieved a pooled AUROC of $0.926\pm0.013$. NoRDeC achieved the highest AUROC in six of seven abnormality categories and in the pooled analysis, while CKA suggested subtype-dependent changes in inter-layer relationships not captured by anomaly scores alone. The normal-reference model is fitted using only normal glomeruli; abnormality labels are used for configuration selection, evaluation, and grouping in the representation analysis. These results show that a frozen renal feature extractor can support both detection and representation-level characterization of glomerular abnormalities without using abnormal examples to fit the detector.
Synthetic data adaptation improves few-shot cryo-ET classification accuracy
Bridging the Synthetic-to-Real Gap for Few-Shot Cryo-ET Classification
Abstract: Subtomogram classification in cryo-electron tomography (cryo-ET) is a challenging problem due to the scarcity of labeled examples. While cryo-ET simulators can be adopted to generate unlimited synthetic data, the substantial domain gap between synthetic and real subtomograms hinders its practical utilization. In this work, we propose a novel synthetic-to-real adaptation framework with a learnable transformation module, bridging this gap at both the input and feature levels. Extensive experiments demonstrate that our method consistently outperforms existing transfer learning baselines in few-shot settings.
Deep learning improves brain surface labeling with limited expert data
Geometric-to-Semantic Spherical Transfer Learning for Cortical Sulci Labeling
Abstract: Deep learning on cortical surfaces faces a dilemma: capturing the complex topology of over 60 nomenclature-dependent sulci per hemisphere requires high-capacity models, yet the extreme scarcity of expert annotations ($N=62$ subjects) inevitably causes overfitting. Standard supervised approaches fail to generalize in this data-scarce regime, particularly for variable and small sulci where topological ambiguity is high. To overcome this limitation, we introduce a Geometric-to-Semantic Spherical Transfer Learning framework. First, we leverage massive unlabeled data (UK Biobank, $\approx$30,000 subjects) to pre-train a spherical encoder using a locally-optimized strategy. By relying solely on continuous surface features (curvature and depth), the relevance of this pre-training is confirmed by the model's ability to detect localized and rare topological traits, such as sulcal interruptions. The downstream labeling task, however, introduces extracted sulcal fundi (lines) as an explicit semantic input. To bridge this dimensional domain gap (from purely geometric to semantic) without causing catastrophic forgetting, these anatomical lines are integrated into the pre-trained backbone via a soft-initialized Topological Prior Injector. Our experiments demonstrate that this approach outperforms fully supervised baselines trained from scratch, achieving a mean Dice of 0.77. Crucially, a local analysis reveals that the self-supervised geometric priors yield the largest performance gains on variable and tertiary sulci (up to 14.8%), confirming that learning the cortex shape is highly beneficial for identifying its rarest parts.
SkNeXt reduces cost of brain cell mapping from huge microscopy data
SkNeXt enables topology-guided neuronal reconstruction from petabyte-scale microscopy data
Abstract: Recent advances in high-resolution fluorescence and electron microscopy have enabled nanoscale imaging across increasingly large brain volumes, but the resulting terabyte- to petabyte-scale datasets make complete neuronal reconstruction prohibitively expensive in computation, data movement, and manual proofreading. Here, we present SkNeXt, a topology-first framework for scalable neuronal reconstruction from large volumetric microscopy datasets. Instead of densely processing entire image volumes, SkNeXt first converts neuronal morphology into compact SWC skeletons that preserve long-range connectivity. Proofreading is therefore focused on sparse neuronal trees, allowing branch, continuity, and connectivity errors to be corrected before high-resolution reconstruction. The corrected skeletons then serve as persistent structural priors for recovering detailed morphology while preserving neuronal identity and topology. Crucially, SkNeXt also uses neuronal skeletons as spatial indices for selective data access, retrieving high-resolution image regions only along reconstructed trajectories and bypassing most background and signal-free volumes. This substantially reduces I/O and computational overhead, allowing reconstruction cost to scale with neuronal morphology rather than total dataset size. Using SkNeXt, we reconstructed neurons from a petabyte-scale super-resolution fluorescence dataset of the mouse brain on a single GPU within one week, without requiring exhaustive dense inference across the complete imaging volume.
Attention guidance improves reliability in multiple instance learning tasks
CAR-MIL: Counterfactual Attention Regularization for Multiple Instance Learning
Abstract: Multiple Instance Learning (MIL) is widely used for weakly supervised learning, particularly in digital pathology, where fine-grained annotations are costly. Most MIL methods aggregate instance features via attention mechanisms. However, attention weights do not always faithfully reflect instance importance and may focus on spuriously correlated regions. In this work, we propose CAR-MIL, a framework that explicitly guides attention learning through a counterfactual attention regularization objective inspired by counterfactual explanations. Built on a standard attention-based MIL architecture, our approach introduces a lightweight counterfactual attention branch trained to produce an alternative prediction while remaining close to the factual attention distribution. This encourages prediction changes to arise from minimal, structured redistributions of attention, leading to more informative evidence allocation. The resulting factual and counterfactual attention maps capture complementary evidence: the former highlights regions supporting the prediction, while the latter reveals regions whose reweighting would challenge it. We evaluate our method on synthetic MIL benchmarks with instance-level ground truth enabling controlled analysis of attention behavior and on five digital pathology datasets across four tasks. CAR-MIL maintains competitive classification performance, with the largest gains observed on more challenging tasks, while improving attention reliability, demonstrating the benefits of integrating counterfactual explainability reasoning into attention learning. Code is available at: https://github.com/ImaneCR/CAR-MIL/.