Papers for

clinical data engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Physiological model test reveals and removes hidden shortcuts in health data

PhysioTRACE: Provenance-Aware Stress Tests for Physiological Foundation Models

Abstract: Physiological foundation models encode how a signal was recorded alongside the physiology it reflects. When recording conditions are associated with diagnosis, this acquisition provenance can become a shortcut, yet the usual evidence, shifted transfer and provenance decodability, does not show whether a predictor uses it. We introduce PhysioTRACE, a four-axis behavioral audit for frozen encoders that separates what a probe can decode from what a fixed task head relies on. Recover scores how decodable provenance is; Stress reverses only the provenance-target association on the same held-out records; Intervene removes a train-localized provenance component; and Verify certifies that removal only if it beats matched random projections within a declared utility margin. Each audit thus ends in one of three verdicts: no reliance, or reliance with the remedy certified or refused. Across EEG and ECG, five training objectives, and five frozen foundation models, the relation between Recover's calibrated score and out-of-distribution utility changes sign between datasets, so neither can stand in for a reliance test. On paired EEG views where the shortcut is known by construction, the audit detects it (the exposed head loses about 0.2 AUROC when the association is reversed, while a control head is unaffected) and certifies removal of a rank-two component that restores control-level behavior without measurable utility loss, for both encoder objectives tested. On real ECG device metadata it returns all three verdicts: it certifies a remedy that removes 91% of one model's excess vulnerability, finds no reliance where device and diagnosis are barely associated, and refuses the remedy for a second model whose localized direction also carries task signal. Robustness to how inputs were recorded therefore needs a behavioral test, and PhysioTRACE provides one that can pass, fail, or refuse a remedy.

Mon 28 SeptMachine Learning
The gist
When medical data is recorded using different devices or setups, models that analyze this data might learn to rely on how the data was collected rather than the actual health signals. The authors created PhysioTRACE, a method to detect when these models depend on such shortcuts and to remove them without hurting model performance. They tested this on brain and heart signal data and showed that PhysioTRACE can correctly find and fix these hidden biases or decide when the fix would harm the model’s usefulness. This helps ensure medical AI systems focus on real health information instead of quirks from how the data was recorded.
Open → 2609.34466v1

Model predicts patient breathing support effects during critical care

Learning response-aware patient dynamics for respiratory support

Abstract: Respiratory support can shape the short-term physiological trajectory of critically ill patients, but patients receiving the same intervention may follow different physiological trajectories. Clinical patient dynamics models typically predict future states from recent physiology and recorded interventions, while physiological change is mainly represented through the predicted future state. We propose a response-aware patient dynamics model that explicitly represents physiological change during autoregressive state updating. The model decomposes predicted physiological change into state-dependent baseline dynamics and respiratory-support-associated deviations, with room air providing a reference for the decomposition. We provide a formal analysis of this reference-anchored formulation. A response pathway encodes the predicted physiological change and uses it to update the latent patient state across the forecast horizon. Across ICU cohorts from two independent institutions, the proposed model achieves comparable overall trajectory prediction to patient dynamics baselines, with more consistent improvements when physiological states are changing.

Sat 26 SeptArtificial Intelligence
The gist
Breathing support helps critically ill patients, but different patients respond differently over time. The authors created a model that better understands how patients’ functions change with breathing support, by separating normal changes from those caused by the support. This helps predict how a patient’s condition might evolve, especially when their health is actively changing. The model was tested on data from two hospitals and showed improved predictions in these situations.
Open → 2609.32782v1

GeoRVQ improves accuracy and quality of physiological signal tokens

GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals

Abstract: Residual vector quantization (RVQ) turns physiological waveforms into compact token sequences, but conventional masked modeling treats every incorrect token as equally costly. We propose GeoRVQ, a coarse-to-fine masked token model whose objective reflects the local response of a frozen waveform decoder. Decoder-induced costs define geometry-aware soft targets and expected distortion, while quantizer-causal prediction follows residual dependencies from coarse to fine levels. In a descriptive aggregate over MIMIC-IV Waveform, VitalDB, and CODE-15\%, GeoRVQ increases exact token accuracy from $.133\pm.004$ to $.143\pm.003$, reduces decoded distance from $.606\pm.006$ to $.393\pm.007$, and increases R-peak F1 from $.784\pm.004$ to $.837\pm.008$ under matched model and training conditions. Across 45 held-out code substitutions, decoder-induced cost has a Spearman correlation of $.85$ with realized decoded cost, compared with $.54$ for Euclidean codeword distance. These results indicate that decoder-aware objectives can improve waveform and event preservation without requiring a large increase in exact token accuracy.

Tue 22 SeptMachine Learning
The gist
Turning heart and other body signals into short sequences helps computers analyze them faster. The authors designed a new method called GeoRVQ that judges errors based on the actual impact on signal quality, not just whether tokens match perfectly. Their approach improves accuracy and better preserves important features like heartbeat peaks. They tested GeoRVQ on major medical datasets and showed it predicts signal details more faithfully than previous methods.
Open → 2609.27018v1

Resnet u-net decoder improves ecg wave boundary detection accuracy

Decoder Design Matters for ECG Delineation

Abstract: Electrocardiogram (ECG) delineation identifies the boundaries of P waves, QRS complexes, and T waves, providing structural annotations that can guide AI models in learning to interpret ECGs. However, training accurate delineation models requires manual annotations that are scarce and time-consuming to obtain. Recent work addresses this limitation through semi-supervised learning (SSL), but the design of the architecture, particularly the decoder, has received less attention. To this end, we propose R-U-Net, an ECG delineation model that pairs a ResNet-18 encoder with a U-Net decoder. On SemiSegECG, R-U-Net outperforms the strongest evaluated ResNet-18 + fully convolutional network (FCN) head baseline in each of the 16 in-domain settings by 3.3-13.0 mIoU and achieves 82.6 mIoU in the cross-domain setting, an improvement of 8.1 mIoU. Controlled ablations show that decoder design contributes more to performance gains than the evaluated SSL methods, motivating further exploration of architectures for ECG delineation. All code is open-source at github.com/ELM-Research/ECG-Delineation.

Tue 15 SeptMachine LearningArtificial Intelligence
The gist
Identifying the start and end of important heart signals in an ECG helps computers better interpret heart health. Usually, training such models takes a lot of expert-labeled data, which is hard to get. The authors show that using a specific combination of neural network parts—called ResNet-18 as the encoder and U-Net as the decoder—improves accuracy more than previous methods. Their work also highlights that how the decoding part of the model is designed matters more than the semi-supervised learning techniques used before.
Open → 2609.16489v1

Alzheimer disease speech detection improved across different settings

Robust Cross-Domain Speech-Based Alzheimer's Disease Detection via Iterative Adversarial Self-Training

Abstract: As Alzheimer's disease (AD) has increasingly become a major global public health issue, speech-based AD detection has attracted widespread attention. However, most existing methods are trained and evaluated on a single dataset, often leading to severe cross-domain performance degradation due to reliance on dataset-specific artifacts rather than disease-related speech cues. In real-world applications, reliable Alzheimer's disease detection requires models that are robust to variations in recording environments, speakers and data collection conditions. To address this challenge, this paper adopts unsupervised domain adaptation to learn robust, domain-invariant feature representations in the absence of target-domain diagnosis labels. On this basis, a novel unsupervised domain adaptation method, Iterative Adversarial Self-Training (IAST), is proposed. Results demonstrate that IAST significantly improves the generalization ability and robustness under various cross-domain settings.

Sat 12 SeptSound
The gist
Detecting Alzheimer's disease by analyzing speech is helpful but challenging when methods only work well on one specific dataset. The authors designed a new approach called Iterative Adversarial Self-Training (IAST) that helps the detection system work reliably even when speech data come from different places or conditions. They do this by teaching the system to focus on disease-related speech patterns and ignore unrelated differences between datasets. Their method improves accuracy when tested across various recording environments and speaker groups.
Open → 2609.14139v1