Papers for

medical imaging analysts

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multimodal model splits task signals from noise for better prediction

Structured Latent Modeling for Supervised Multimodal Information Decomposition

Abstract: Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or masked objectives at intermediate layers, coupled with source-wise invertible normalizing flows and a supervised, low-rank latent variable model. This architecture explicitly factorizes the joint distribution into shared task-relevant variation, modality-specific predictive variation, and task-irrelevant dependence. Drawing connections to prior multimodal learning assumptions, our approach evaluates how modalities independently and jointly contribute to the target. Ultimately, this framework unites intermediate representation learning with structured likelihood-based guidance, offering a practical latent-variable lens for characterizing continuous multimodal interactions. Empirically, we demonstrate the effectiveness of our approach across diverse multimodal benchmarks, showing robust improvements in predictive performance.

Mon 28 SeptMachine Learning
The gist
When computers look at multiple types of input, like images and text together, they need to figure out which parts of these inputs really help solve a problem and which parts are just noise or unrelated. The authors created a new method that breaks down the information into parts that are shared across all inputs, parts unique to each input type, and irrelevant parts, all in a structured way within the computer’s internal reasoning. This helps the computer understand how each kind of input adds to solving the problem, leading to better predictions. They showed this works well on various tests with different types of data.
Open → 2609.35502v1

Model improves visual answer accuracy by learning from decision boundaries

BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation

Abstract: Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional observations or inter-sample comparisons, yet a visually valid cue is not necessarily decision-relevant: it may describe how samples differ without changing the model's relative preference between competing answers. We therefore introduce BIRD, a self-improving Boundary-Informed Rationale Distillation framework that uses model-specific confusions to locate unresolved local decision boundaries and distills the evidence that resolves these confusions into rationales. For each sample, BIRD retrieves candidate neighbors from the target MLLM's own representation space and selects the most confusable one according to its answer preferences. It then generates answer-blind candidate evidence from their visual differences and functionally verifies which evidence most effectively strengthens the model's preference for the correct answer while avoiding inappropriate transfer across the pair. The verified evidence is then distilled into a single-sample rationale for standard supervised fine-tuning. Experiments on medical and chart VQA show that BIRD outperforms competing rationale-augmentation methods across two target MLLMs, while further analyses demonstrate clearer separation of confusable answers and stronger gains from model-matched supervision.

Sun 27 SeptArtificial Intelligence
The gist
It can be hard for AI that looks at pictures and answers questions to tell closely related answers apart, especially in special fields like medicine. The researchers designed a method called BIRD that finds tricky cases where the AI is unsure and learns exactly which visual clues help it choose the right answer. This approach teaches the AI with clearer explanations based on those clues, helping it make better decisions. Testing showed this method works better than similar ways to improve AI understanding in areas like medical images and charts.
Open → 2609.33713v1

Brain signal analysis improves by modeling complex noise patterns

Beyond Gaussian Assumptions: Distribution-Aware Channel Capacity for Effective Connectivity

Abstract: Effective-connectivity estimation from brain signals often relies on Gaussian residual modeling, which enables tractable estimation but can discard informative distributional structure and distort inferred directed interactions when empirical residuals are non-Gaussian. We show across multiple modalities, species, and experimental conditions that both brain signals and fitted channel residuals frequently deviate from Gaussianity. We therefore introduce a distribution-aware, information-theoretic measure of effective connectivity based on channel capacity under general residual distributions. To estimate the resulting capacity from empirical, potentially non-Gaussian residuals, we develop a dual-flow min-max estimator based on normalizing flows, in which a generator searches over admissible input distributions under a power constraint while an observer estimates output entropy. We provide a theoretical characterization of the estimator, showing that the observer objective recovers differential entropy up to a KL approximation term, that the formulation reduces to classical Gaussian capacity as a special case, and that residual entropy can alter achievable information rates beyond variance; game-gap and error analyses further characterize optimization and approximation sources. In brain-like simulations with known directed connectivity, Dual-flow achieves the highest AUROC and AUPRC across ten conditions spanning diverse network topologies, hidden drivers, feedback, and heterogeneous hemodynamics, compared with Gaussian capacity, Granger causality, VAR-LiNGAM, and GIMME. Applied to multimodal brain signals, the method reveals time- and condition-resolved directed interactions consistent with known neurobiological circuitry. Together, these results establish a principled distribution-aware framework for effective-connectivity estimation beyond Gaussian residual modeling.

Sat 26 SeptMachine Learning
The gist
Estimating how different parts of the brain interact often assumes a simple type of noise called Gaussian noise, which can miss important details. The authors found that brain signals and their noise usually don’t follow this simple pattern. They created a new way to measure brain connectivity that considers these more complex noise patterns, resulting in more accurate insight into brain interactions. Their method uses advanced mathematical tools to better capture information flow in brain signals across different species and conditions.
Open → 2609.32774v1

Brain-age prediction improved by modeling stable multifractal patterns

Beyond Feature Reliability: Repeat-Informed Multifractal Curve Regression for Brain-Age Prediction

Abstract: Brain-age prediction from resting-state fMRI provides a quantitative framework for characterizing age-related changes in spontaneous brain dynamics and for identifying functional signatures. Existing studies have linked fractal and multifractal scaling to age and examined the reliability of individual features. However, prediction repeatability depends on how features fluctuate jointly and how a predictor combines them, which feature-wise reliability assessments do not capture. To address this problem, we propose Repeat-informed Multifractal Curve Regression (RMCR), a structured framework for learning stable age-predictive patterns from multifractal curves. By jointly modeling curve structure and repeat-scan variability, RMCR learns predictive combinations of fluctuation orders that target both accuracy and within-subject consistency. Relative to a matched run-level ridge baseline, RMCR reduces single-run MAE by 6.1% on HCP-A and 7.9% on an external Cam-CAN cohort, and within-visit repeat absolute difference by 18.5% on HCP-A, using a single scan at inference.

Thu 24 SeptMachine Learning
The gist
Predicting a person’s brain age using resting brain scans can help understand how brain activity changes with age. The authors introduce a new method called RMCR that looks at complex patterns in brain signals and how these patterns are consistent across repeated scans. This approach not only makes age predictions more accurate but also ensures the predictions are more reliable when scans are repeated. They tested RMCR on two separate brain scan datasets and found it improved prediction accuracy and repeatability compared to existing methods.
Open → 2609.29307v1

Active spot selection does not clearly beat random sampling in spatial transcriptomics

Benchmarking Active Spot Selection for Cost-Efficient Spatial Transcriptomics

Abstract: Spatial transcriptomics (ST) measures gene expression in tissue context, but dense capture grids can be costly and may repeatedly sample morphologically similar regions. Most active learning strategies were developed for categorical labels and independent samples. We conduct a retrospective pool-based benchmark of active learning versus uniform Random sampling for ST, where expression vectors are high-dimensional and continuous and candidates are spatially correlated. Using two fully profiled public ST cohorts, we mask candidate expression vectors and simulate multi-round selection with uncertainty-based Monte Carlo dropout (MC-dropout) and temporal output discrepancy (TOD), and diversity-based CoreSet and TypiClust-inspired selection. We compare 160 completed configurations at 5%, 10%, 30%, and 50% of the fold-wide training spot pool under patient-level cross-validation, with a separate full-label reference. Within each budget, strategies share the selection schedule, morphology-to-expression predictor, and optimization protocol. We assess mean per-gene within-slide Pearson correlation coefficient (PCC), expression-cluster agreement, and Moran's I fidelity. On HER2-positive breast cancer, pooled mean PCC differences from Random across the four active strategies were -0.0176, -0.0117, +0.0056, and +0.0057 at 5%, 10%, 30%, and 50%, respectively. On cutaneous squamous cell carcinoma (cSCC), three strategies were below Random at 5%, and all four were below Random at 10%. On HER2-positive breast cancer, CoreSet and MC-dropout had lower PCC but higher expression-cluster agreement than Random at the two smallest budgets; this pattern did not reproduce on cSCC. Under the reported fixed training horizons, the evaluated active strategies do not consistently improve on Random at small budgets, and rankings depend on the evaluation measure.

Wed 23 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Spatial transcriptomics measures gene activity across different regions in tissue, but it can be expensive to collect data everywhere. The authors tested smarter ways to pick spots to measure so they could save costs. Using two public datasets, they found that these smarter methods did not consistently give better results than just picking spots randomly. The effectiveness depended on how many spots were sampled and the measurement used to judge performance.
Open → 2609.27208v1

A new method for detecting anomalies by modeling their causes

A Principled Approach to Unsupervised Anomaly Detection

Abstract: Traditional unsupervised anomaly detection (UAD) methods are designed to flag or localise deviations from a normative distribution, ignoring the underlying generative mechanisms of the anomalies. Yet the nature of an anomaly is often as important as its presence. We reformulate UAD as a Bayesian inverse problem, in which the objective is to infer the most probable corruption responsible for each observation. Our framework yields a probabilistic anomaly score as the energy of the inferred corruption parameters, and serves as a principled recipe for developing new UAD algorithms. We derive several existing methods as instances of the general framework, each corresponding to the same energy score under different modelling choices. Experimentally, we study the framework's components in a controlled setting, and improve object-class AUROC on the MVTec AD dataset by 2.3% by adapting the underlying corruption model. Finally, we validate the framework on a brain MRI benchmark, achieving strong detection performance while producing estimates of pathology intensity, bias, and geometry. Code is available at https://github.com/jgmyles/inverse-uad.

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
Detecting anomalies usually means spotting anything unusual, but often it’s important to understand what caused those unusual things. The authors propose a new way to detect anomalies by guessing the hidden changes that made the data strange, using probabilities. Their method can explain how bad changes happened, not just if something is wrong. This approach improved detection accuracy in industrial images and brain MRIs, while also estimating how the abnormalities appeared.
Open → 2609.21800v1

Scientific image quality assessed by multimodal retrieval system

Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

Abstract: This paper proposes a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, designed to simultaneously address both the understanding track (SIQA-U) and the scoring track (SIQA-S) of the SIQA challenge. We construct a multimodal index that integrates textual semantics with fine-grained visual features, and develop a multi-route retrieval and fusion mechanism to provide large language models with highly relevant reference cases, thereby enhancing their capability to evaluate complex scientific images. Experimental results demonstrate that the proposed framework effectively aligns with the judgment criteria of human experts. Ultimately, our method achieves 1st place in the SIQA-U track of the SIQA challenge at the ICME 2026 Grand Challenges.

Thu 17 SeptComputer Vision and Pattern RecognitionComputation and Language
The gist
Judging the quality of scientific images can be tricky because it requires understanding both the pictures and their scientific context. The authors created a system that helps a language model by giving it relevant visual and text examples to compare with the images being evaluated. This approach helps the system better think like human experts when assessing image quality. Their method won first place in a scientific image quality assessment challenge.
Open → 2609.19634v1

Time-warping estimation improves signal analysis with faster computing

Time-warping estimation via stationarity-based learning of the de-warped signal

Abstract: Time-warping estimation is a fundamental problem in signal processing with applications in bioacoustics, radar, and biomedical analysis. This paper introduces a Time-Warping Estimation Trainable (TWET) model for estimating timewarping functions from a single observation. The proposed approach formulates time-warping estimation as a stationarization problem in the wavelet domain and leverages a hierarchical dilated convolutional architecture to estimate the time-warping functions. A differentiable stationarity criterion is introduced for end-to-end optimization. TWET is compared with existing approaches. Experimental results show improved deformation reconstruction accuracy together with significantly reduced computation time, making the framework compatible with low-latency applications.

Tue 15 SeptMachine Learning
The gist
Time-warping is like adjusting the speed of a recording to line up important events. The authors created a new method called TWET that uses a smart wavelet approach to guess how a signal has been stretched or squished in time from just one example. This method teaches a computer model to find these changes quickly and accurately by looking for stability patterns in the signal. Compared to older methods, TWET is faster and better at fixing these time changes, making it useful for things like medical or animal sound analysis.
Open → 2609.16796v1

PSMP-CLIP improves zero-shot anomaly detection with better image masks

PSMP-CLIP: Patch-Prompt SAM and Multi-Semantic Prompting for CLIP-Based Zero-Shot Anomaly Detection

Abstract: Zero-shot anomaly detection aims to localize anomalies without target-domain samples. Existing CLIP-based methods suffer from coarse anomaly maps and limited semantic prompts. We propose PSMP-CLIP, integrating patch-prompt SAM2 segmentation (PPSS) and multi-semantic guided prompt regularization (MSGPR). PPSS samples prompts directly from intermediate patch features, avoiding threshold drift and guiding SAM2 to produce precise masks. MSGPR uses multiple learnable prompts constrained by semantic anchors to preserve generalization. Experiments on 14 datasets show highly competitive performance, achieving the best pixel-level AUROC on MVTec AD, BTAD, DTD-Synthetic, CVC-ClinicDB, TN3K, Endo, and Kvasir.

Tue 15 SeptComputer Vision and Pattern Recognition
The gist
Detecting unusual parts in images without using examples from that image type is hard. Existing methods using a model called CLIP struggle to pinpoint odd spots precisely. The authors present PSMP-CLIP, which helps by creating better image masks from smaller image patches and by using multiple guided text prompts for improved detection. They tested this approach on many datasets and found it works better than similar methods in identifying anomalies at the pixel level.
Open → 2609.16785v1

AstroSpecLM links astronomical spectra with language model answers

AstroSpecLM: A Spectrum-Language Model for Evidence-Grounded Astronomical Spectral Analysis

Abstract: Astronomical spectra encode rich physical information, but drawing scientific conclusions from spectral features typically requires expert interpretation. This paper presents AstroSpecLM, a spectrum-language model that connects one-dimensional DESI spectra with Qwen3-4B to answer questions and provide explanations grounded in spectral evidence. Instead of generating question-answer pairs directly from templates or raw catalog fields, we first distill each spectrum into a compact set of catalog- and spectrum-derived facts, then use these facts as references to generate instruction-following conversations. The resulting model is competitive with specialist supervised baselines on classification and redshift estimation, while additionally producing natural-language explanations that reference specific spectral features. Our results indicate that grounding a language model in one-dimensional scientific spectra is feasible, and that fact-mediated instruction data yields a model capable of both prediction and explanation.

Mon 7 SeptArtificial Intelligence
The gist
Astronomical spectra contain detailed clues about celestial objects, but understanding them usually needs expert knowledge. The authors created AstroSpecLM, a model that reads these spectra and answers questions in natural language, explaining its reasoning based on specific features it sees. They trained it to first extract important facts from the spectral data before generating answers and explanations. This approach performs well compared to specialized methods and makes the analysis more accessible.
Open → 2609.07102v1