Papers for

hospital data teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

UID preserving method improves clinical event timelines from discharge summaries

Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication

Abstract: Clinical timelines support treatment-window analysis and leakage-free modeling, but discharge summaries often obscure chronology and structured EHR tables describe only part of the patient course. We present a UID-preserving framework that links each narrative event occurrence to its source span and retains that identity through text-only estimation, structured-evidence retrieval, timestamped source-row grounding, and joint revision. We also present GAVEL, an LLM judge that compares two UID-aligned timelines against the narrative and structured record, to augment prior matching and temporal assessments. Across six open-weight models and 40 mixed-critical-care summaries, the GLM 5.2 multimodal revision, as compared to its text-only variant, improved temporal agreement without reducing event recovery and performed competitively with clinician annotations, while other model revisions showed smaller gains and lower overall performance. Ablations showed that UIDs primarily preserve event retention, whereas source-row linkage supports temporal placement. Blinded human review upheld most GAVEL findings, and controlled adjudication favored multimodal over text-only GLM 5.2 but did not for DeepSeek V3.2. In developing the UID and judge pipeline, we are able to demonstrate 43\% increased event recovery, a framework competitive with clinician annotations, and a system with occurrence-level provenance for both reconstruction and evaluation.

Fri 11 SeptArtificial Intelligence
The gist
Medical records often list events out of order or miss details, making it hard to understand a patient's treatment timeline. The authors created a system that links each event in a doctor's notes back to its source and uses both text and record data to build accurate timelines. They also built a tool that checks and compares these timelines against the original records. Their approach recovered many more events and matched expert doctors’ timeline assessments closely.
Open 2609.13062v1

Parallel training improves covid diagnosis model speed and accuracy

Parallel Training Using a CNN-DNN Architecture for Accelerated Development of Diagnostic Models

Abstract: Artificial intelligence has shown promise in assisting radiologists in imaging-based diagnosis across a wide range of diseases. Efficient training of large deep learning models is essential to cope with extremely large data sets or dynamically growing disease data, like in a pandemic like situation. In this retrospective study, we collected 300 CT scans from COVID-19 and non-COVID-19 pneumonia patients from three different centers in Germany. We investigated a hybrid CNN-DNN network model based on image decomposition and localization that naturally supports parallel and efficient training of deep learning models. In total, 156 models with three different architectures were trained to capture features at different levels resulting in 12 patient-level COVID-19 diagnosis models. Diagnostic performance as well as time saving were measured. The highest accuracy was obtained from DenseNet121 and 3D CNN models with a parallel CNN-DNN approach, resulting in $88.78\%$ training, $76.67\%$ validation and $76.03\%$ test accuracy for the DenseNet121 with $4\times4\times1$ subdomains and $87.72\%$ training, $76.82\%$ validation and $74.86\%$ test accuracy, respectively, for the 3D CNN with $4\times4\times1$ subdomains. The strongest reduction in parallel training time by a factor of $31$ was observed for the 3D CNN model and $4\times4\times2$ subdomains. Our parallel training approach improves efficiency as well as performance enabling rapid model development, among others crucial for pandemic preparedness.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Training deep learning models to detect diseases in medical images can take a long time, especially with large and growing datasets like during a pandemic. The authors collected CT scans from COVID-19 and other pneumonia patients and used a combined CNN-DNN approach that breaks images into parts to speed up training by running in parallel. This method achieved similar or better accuracy while reducing training time significantly. Their approach helps create diagnostic tools faster, which can be important in urgent health situations.
Open 2609.12902v1

PrecepTron enables large scale evaluation of medical AI reasoning

Scaling Clinical Judgment to Evaluate Medical AI

Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also release GRAND-ROUNDS, a new large-scale physician-annotated benchmark of 9,217 scores by 11 physicians across seven studies. We show that frontier LLMs in typical "LLM-as-a-judge" approaches often disagree with physicians and with each other, but fine-tuning PrecepTron on a small number of cases enables physician-level consistent scoring across tasks. We use PrecepTron to reproduce headline findings from five influential studies assessing LLMs for clinical care in JAMA, Science, and Nature Medicine without new human grading. Using PrecepTron, we then pose new questions about how LLMs reason in medicine that would have been infeasible with human grading alone, including measuring the diagnostic accuracy of frontier LLMs when clinical cases are provided piecemeal, even token by token. Together, PrecepTron and GRAND-ROUNDS provide a foundation for reproducible, large-scale study of how LLMs reason in medicine. All code, data, and labels are made freely available for researchers.

Fri 11 SeptArtificial Intelligence
The gist
Evaluating how well AI models think like doctors is hard because it usually needs many doctors reviewing answers, which is expensive and slow. The authors created PrecepTron, an AI trained to judge medical AI answers like a real doctor, using only a small set of examples from physicians. They also made a big set of doctor-scored questions called GRAND-ROUNDS. With PrecepTron, they could check how medical AI models reason without needing doctors to grade everything, and even explore new ways to test AI on medical cases bit by bit.
Open 2609.12822v1

Meddies improves clinical data privacy with multilingual PII detection

Meddies-PII: A Multilingual Framework for Personally Identifiable Information Extraction in Clinical De-identification

Abstract: Clinical de-identification relies on accurately identifying personally identifiable information (PII). However, manually annotated datasets are costly to construct, while existing synthetic alternatives often provide limited details about their generation process or rely on relatively simple synthesis strategies. We introduce Meddies-PII-Dataset, a corpus of one million synthetic clinical documents spanning seventeen languages and nine PII labels. The documents are generated using attribute-conditioned prompts and validated through thirteen deterministic gates that enforce structural and annotation consistency. To evaluate the dataset's utility, we train Meddies-PII-Model, a BIOES token classifier, and compare it with existing PII extraction systems using exact-match entity-level F1. Meddies-PII-Model achieves the highest performance among the evaluated systems on all reported benchmarks, with a mean F1 of 0.827 across fifteen external benchmarks, compared with 0.658 for the strongest baseline. Upon acceptance, we will publicly release the dataset, benchmark suite, model, generation framework, and evaluation code to support research on multilingual clinical de-identification.

Fri 11 SeptComputation and LanguageArtificial Intelligence
The gist
Protecting patient privacy in medical records requires identifying personal information correctly. Creating real training data for this is expensive, and previous synthetic alternatives were less detailed. To solve this, the authors made a huge set of one million fake clinical documents in 17 languages, carefully checked for accuracy. They trained a model to find personal details in these documents and found it outperformed other methods across many test sets. They will share their dataset, model, and tools publicly for others to use.
Open 2609.12544v1

Deep learning combines brain scans and clinical data to diagnose alzheimer’s

A Multimodal Explainable Deep Learning Framework for Alzheimer's Disease Diagnosis using 3D Magnetic Resonance Imaging and Clinical Data

Abstract: Dementia is a major and growing global health burden, with Alzheimer's disease (AD) accounting for most cases. Timely and accurate diagnosis is central to managing this burden and increasingly depends on integrating complementary clinical and imaging information. Multimodal deep learning can combine these modalities for AD diagnosis, but how its explanations behave across modalities, fusion strategies, and cohorts remains unclear. We developed an explainable multimodal framework pairing a 3D CNN encoder for T1-weighted MRI with a feedforward network for harmonized clinical and demographic data, comparing varied model setups on three-way and pairwise diagnostic tasks using 6,479 internal records from the ADNI and 1,703 independent records from the OASIS-3. On ADNI, the tabular-only model achieved the highest three-class AUC-ROC of 0.879 and best discriminated cognitively normal (CN) versus mild cognitive impairment (MCI; 0.903), while cross-attention performed best for MCI versus AD (0.861); CN versus AD was highly discriminative overall. On OASIS-3, the vision-only model performed best (three-class AUC-ROC 0.910); CN versus MCI remained difficult, and no fusion strategy consistently outperformed single modalities across tasks and cohorts. SHAP and Integrated Gradients identified the MMSE as the dominant tabular feature in both cohorts, with global feature rankings agreeing strongly in ADNI ($ρ=0.94$) and OASIS-3 ($ρ=0.96$); CAM-based explanations, however, changed with model configuration and cohort. These findings show that multimodal performance and explanations are task, modality, fusion, and cohort-dependent: a dominant cognitive signal persisted across cohorts, but feature contributions and CAM explanations did not, underscoring the need to evaluate explainability under cohort shift rather than as a stable, intrinsic property.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Diagnosing Alzheimer's disease early is important but challenging, and doctors use brain scans and clinical tests to help. The authors created a computer model that learns from both 3D brain images and patient information to improve diagnosis. They found that the model’s accuracy and explanations vary depending on the data type, how the information is combined, and the patient group. A key mental test consistently helped diagnosis across different groups, but the model’s attention to brain scan features changed. This highlights the need to check how explanations hold up when testing on new patients.
Open 2609.12410v1

AI system reduces false ICU alarms while limiting missed alerts

Certified AI Triage of ICU Alarms

Abstract: In the VTaC benchmark 71% of ventricular-tachycardia alarms are false, but silencing a real one can delay recognition of a dangerous arrhythmia. We reframe alarm reduction as three-way triage (retain, suppress, or defer) and bound the decision this analysis treats as harmful: among suppressed alarms, the fraction that were genuine stays below a user-set budget with 95% confidence, under i.i.d. event sampling. Alarms sharing a waveform record are dependent, so the clustered analysis is a sensitivity check. On the official split a 5% budget certifies in all three seeds, suppressing 74.8% of false alarms while silencing 1.5% of genuine ones, at AUROC 0.953 and Challenge Score 83.33, numerically comparable to the strongest of the eleven published systems. Our central finding measures what multiplicity costs: the correction charges for every candidate, so a finer grid can certify strictly less. Under held-out calibration the 885-cell grid we declared certifies 1 of 15 fold-runs, while choosing the grid on a separate selection partition certifies 8. We project the calibration volume each budget needs, making an uncertifiable budget a design parameter. Finally, adding a learned reliability dimension to the policy grid did not sharpen the certified frontier.

Fri 11 SeptMachine Learning
The gist
Too many alarms in intensive care units can be false, making it hard for staff to notice real emergencies quickly. The authors developed a method that either keeps, ignores, or delays alarms to reduce false alerts while ensuring very few real emergencies are missed. They provide a mathematical guarantee that the number of missed real alarms stays below a chosen limit with high confidence. Their method performs well compared to other systems on a standard test, reducing false alarms by nearly 75% while only silencing about 1.5% of genuine ones.
Open 2609.12365v1

Reinforcement learning improves clinical reasoning in ehr models

Reinforcement Learning over Patient Trajectories for Clinical Reasoning in EHR Foundation Models

Abstract: Electronic health record (EHR) foundation models trained on longitudinal patient trajectories have demonstrated strong performance across diverse clinical prediction tasks. However, their clinical reasoning capabilities remain constrained by next-token prediction on limited and incomplete EHR data. To address this, we propose a reinforcement learning (RL) fine-tuning framework that treats EHR foundation models as generative policies over patient trajectories. We formulate common clinical prediction problems (e.g., hospital readmission) as event-conditioned, time-windowed reasoning tasks. We then design time-aware, rollout-sensitive rewards to account for finite rollout lengths and temporally inconclusive outcomes. We find that RL fine-tuning consistently improves over pre-trained backbones and strong baselines. Notably, it enables smaller models to surpass larger pre-trained models in data-limited regimes and induces positive transfer across tasks. Further analysis shows that RL fine-tuned models generate trajectories with stronger structural and semantic alignment to ground truth and greater downstream utility.

Thu 10 SeptMachine LearningArtificial IntelligenceComputers and Society
The gist
Electronic health records store a lot of patient information over time, but computer models that analyze this data often struggle to think through medical problems carefully. The authors improved these models by training them with a reinforcement learning method that helps the models consider sequences of patient events more thoughtfully. This approach made the models better at predicting outcomes like hospital readmission, even with less data and smaller models. The improved models also create patient event sequences that match real data more closely and are more useful for other clinical tasks.
Open 2609.12277v1

Differential privacy protects patient data in clinical EEG features

Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical Neurophysiology

Abstract: Clinical electroencephalography (EEG) data are valuable for healthcare research and for developing artificial intelligence (AI)-based clinical decision-support systems, but EEG recordings and derived features may contain sensitive patient-specific information. This creates privacy risks when data are reused, analyzed, or shared across clinical and research environments. Conventional anonymization methods are often insufficient for high-dimensional biomedical signals, since removing direct identifiers does not necessarily prevent re-identification, linkage, or inference risks. At the same time, strong privacy protection may distort clinically relevant signal characteristics and reduce data utility. This paper studies subject-level differential privacy for protecting clinical EEG-derived feature representations using Gaussian and Laplace perturbations. The proposed framework considers three deployment scenarios: client-side anonymization, centralized server-side anonymization, and decentralized local training. Following EEG preprocessing and feature extraction, Gaussian and Laplace perturbations are applied to the resulting patient-level EEG feature representations. The Laplace experiments evaluate the implemented noise scales, while the scales required for formal full-vector calibration are derived separately. The effects of both perturbations are assessed using statistical utility measures and a downstream machine-learning-based utility check. The results show that differentially private perturbation can be integrated into EEG processing workflows, but the selected mechanism, privacy parameters, and sensitivity calibration strongly influence data utility. The study highlights the practical privacy-utility trade-off in DP-based EEG feature anonymization and the challenges of preserving downstream utility in small and imbalanced clinical EEG datasets.

Thu 10 SeptCryptography and SecurityMachine Learning
The gist
EEG brain data can reveal private information about patients, so keeping it safe when shared or analyzed is important. The authors studied ways to add controlled noise to EEG features to protect patient identity while still keeping the data useful for medical analysis. They tested different methods to see which balance keeping data private and still good for AI tools. Their results show protecting privacy is possible but tricky, especially with small or uneven patient data. This work helps understand how to better share sensitive brain data safely.
Open 2609.11777v1

Cross-lingual clinical annotation improves medical text tagging accuracy

Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study

Abstract: Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical corpus construction, and to characterize its robustness and computational trade-offs relative to candidate-based projection pipelines. Methods: We developed a constrained LLM projection workflow that inserts entity tags directly into immutable target-language text, followed by deterministic validation and character-offset reconstruction. We evaluated it alongside supervised candidate-span projection and hybrid ML-LLM refinement for transferring Spanish Disease, Symptom, and Procedure annotations into six languages. Evaluation used MultiClinAI gold standard with strict span matching and character-overlap F1 Results: Direct LLM projection achieved the strongest and most consistent performance. GLM 5.2 obtained a mean Strict F1 of 0.9201 across 18 language-entity combinations, while locally deployable Gemma4:31B achieved 0.9133. The best LLM configuration improved Strict F1 over the previous state of the art in all 18 settings, by 0.0564-0.1512, yielding 55,416 grounded mentions with reconstructed offsets. Conclusions: Direct LLM-based projection enables high-quality multilingual clinical annotation transfer and provides a practical approach for extending clinical NLP resources to languages with fewer annotated datasets and language-specific tools. Combined with local inference and deterministic validation, it can substantially reduce expert time and cost for multilingual clinical corpus construction.

Thu 10 SeptComputation and LanguageArtificial Intelligence
The gist
Labeling medical text in many languages is hard because each language needs special tools and data. The authors tested a new way to use large language models (LLMs) to add medical labels directly into text in six different languages. This approach worked better than older methods that picked label spans in the text. Their method helps make accurate and reusable medical datasets across languages with less expert work and cost.
Open 2609.11450v1

Emergency department revisit screening improves with AI knowledge graph

Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support

Abstract: Background: Emergency Department (ED) return visits are commonly reviewed for quality assurance, but are often limited (e.g., to revisits within 48-72 hours) to increase actionable finding yield while minimizing chart review burden. Those limitations may lead to missed quality improvement opportunities. Methods: We conducted an exploratory, retrospective study of randomly selected ED visits to a multihospital health system having an ED revisit within 1-14 days to the same health system. Given only each visit's primary diagnosis, raters (2-3 clinicians and GPT-4 large language model [LLM]) assessed characteristics of the diagnosis pairs, including the "target": whether a pair warranted further assessment. Informed by rater response analyses, an algorithm leveraging an LLM-populated knowledge graph ("KGA") was created to automatically screen for potentially concerning pairs, then preliminarily assessed. Results: 99 diagnosis pairs were included. GPT-4 responses poorly correlated to clinician raters, rating nearly all (94%) pairs as warranting follow-up (4.4-13.3 times more than clinicians). However, prompt engineering was minimal. Among clinician raters, revisit medical gravity was consistently significantly associated with the target, while a differential diagnosis/complication composite was significantly associated on unadjusted, but not adjusted (though less powered) analysis. The KGA achieved 83-100% positive predictive value for at least one clinician rater determining further assessment was warranted based on the diagnosis pair. Conclusion: These results can inform next steps for improving screening with LLMs like ChatGPT. Further research is warranted to validate this preliminary work's finding that the KGA may enable enhancing the scope and yield of screening without substantially increasing reviewer workload.

Wed 9 SeptComputers and SocietyArtificial Intelligence
The gist
Emergency departments check patients who return soon after their first visit to find ways to improve care. The study looked at doctors and an AI called GPT-4 to decide if follow-up reviews were needed based on diagnosis pairs. GPT-4 suggested follow-ups almost all the time, much more than doctors, possibly because it was not specially guided. The researchers built a tool using AI and a knowledge graph that better matched doctors’ judgments, helping flag important cases without adding too much review work. This work points toward smarter AI helping hospitals focus on the most concerning return visits.
Open 2609.10421v1

OmniMed FL fuses images and notes for safer clinical diagnosis

OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis

Abstract: Simultaneous assessment of medical imaging and patient records is often required in clinical diagnosis. However, standard machine learning algorithms cannot analyze these data types together. Meanwhile, compliance with HIPAA and GDPR can constrain centralized aggregation of sensitive patient data. This leaves a crucial void of secure fusion of visual and textual context across distant networks. Thus, we present OmniMed-FL, a controlled systems study of multimodal federated learning for five-class clinical condition classification (Normal, Pneumonia, COVID-19, Pleural Effusion, Cardiomegaly). Our proxy corpus pairs 3,000 public chest radiographs with 3,000 class-conditioned synthetic notes, matched by class, not by patient. The framework benchmarks eight fusion strategies, three initializations, four missing-text imputation rules, and matched federated baselines under non-IID Dirichlet partitioning across 3 to 20 hospital clients. As all notes are synthetic and pairing is not patient-level, these are descriptive proxy comparisons, not estimates of diagnostic performance or deployment readiness. Within those limits with clients ($K=5$) and severe skew ($α=0.1$), local-only training achieves a macro-F1 score of 0.297, FedAvg achieves $0.662\pm0.074$, FedProx $0.737\pm0.085$, a matched FedMME-style one-shot ensemble $0.647\pm0.080$, and our SCAFFOLD-AdamW adaptation $0.070\pm0.015$, the 0.075 FedProx-FedAvg gap falling inside the wider of the two two-seed standard deviations. Over a $4\times3$ grid, label skew costs up to 0.27 F1 whereas a near-sevenfold client increase costs at most 0.10, while bidirectional volume grows linearly to 183.5 GiB at $K=20$. Multimodal fusion leads on both corpora, scoring 0.956 against 0.934 for text and 0.664 for images on the synthetic corpus and 0.906 against 0.880 and 0.737 on the radiograph corpus, for $2.3\times$ the model state of text alone.

Wed 9 SeptMachine LearningArtificial Intelligence
The gist
Doctors often need to look at medical images like chest x-rays and read patient notes together to make a diagnosis, but standard computer programs can’t easily combine these two data types. The authors created OmniMed-FL, a system allowing many hospitals to jointly train diagnostic models without sharing sensitive patient data directly, respecting privacy laws. They tested different ways to combine image and text data and handle missing information using simulated data, showing that combining both types improves diagnosis classification compared to using either alone. This work shows promising methods for secure and effective joint medical data analysis across institutions.
Open 2609.10364v1

OntologyAligner improves mapping of biomedical text to concepts

OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization

Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle distinctions among hierarchically related concepts can obscure concept boundaries. We present OntologyAligner, a three-stage framework that combines ontology-aligned retrieval, large language model candidate reranking, and selective hierarchy-guided refinement. We also construct PhenoNormBench, a unified benchmark comprising 13,390 samples from seven Human Phenotype Ontology datasets. OntologyAligner achieved state-of-the-art performance on HPO normalization, with 88.78% Macro Top-1 Accuracy and 86.75% Micro Top-1 Accuracy, exceeding the strongest baseline by 4.85 and 5.07 percentage points, respectively. Ablation analyses showed complementary contributions from all three stages, and sensitivity analyses demonstrated stability across candidate-set sizes and model backbones. Applications to MONDO, MEDIC, and NCBITaxon further established portability to other ontologies. OntologyAligner offers a generalizable framework for accurate mapping of biomedical text to structured ontology concepts. PhenoNormBench and the code are publicly available at https://github.com/zhelishisongjie/OntologyAligner.

Wed 9 SeptArtificial IntelligenceComputation and Language
The gist
Mapping medical terms from free text to standardized concepts can be tricky because words vary and related concepts are often similar. The authors created OntologyAligner, a three-step method that finds potential matches, ranks them using large language models, and then refines choices by considering concept hierarchies. They tested their method on many examples related to human health traits and showed it works better than previous approaches. The method also works well on other biomedical concept databases, making it broadly useful.
Open 2609.10055v1

MedDeID enables hospitals to remove personal data from clinical notes locally

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

Abstract: Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation with model training, inference, pseudonymisation and evaluation. On an independently annotated, adjudicated 300-note Dutch hospital benchmark, a hospital-trained compact transformer detected 98.9% of identifying text while redacting 0.24% of text outside annotated identifiers; a synthetic-only counterpart detected 96.1%. On 100 primary-care notes, the synthetic-trained model achieved higher recall than the hospital-trained model (90.3% versus 87.0%) and greater robustness to identifier-format perturbations. An English instantiation trained without real text detected 99.7% and 98.9% of annotated identifier characters on two external synthetic benchmarks. These results demonstrate transfer of the workflow to another language, but not clinical English performance. MedDeID provides a route to locally governed de-identification using real or synthetic training data.

Wed 9 SeptComputation and LanguageMachine Learning
The gist
Clinical notes often include personal information that makes it hard to use them for research without risking privacy. The authors created MedDeID, a tool that hospitals can run on their own computers to automatically find and hide personal details in clinical texts. The tool can learn from real or computer-generated notes, works well on Dutch and English data, and keeps sensitive information safe without sending it outside the hospital. This helps hospitals share valuable medical data for AI and research while protecting patient privacy.
Open 2609.10049v1

Cross attention improves heart event prediction from medical claims data

In Medical Claims Data, Enhancing Predictive Performance for Major Adverse Cardiovascular Events Using Cross Attention

Abstract: Medical claims data comprise the financial details, including the expenses and billing information, as well as the clinical information, such as the diagnoses and treatments, of patients visiting medical facilities. Recently, it has been acknowledged that large databases can be constructed from medical claims data for medical research purposes. However, the clinical information within these datasets is often medically unstructured, limiting its application in comprehensive analyses. This study enhances predictive model performance for major adverse cardiovascular events (MACE), a leading cause of death worldwide. Models that predict MACE are crucial to clinical practice guidelines. We utilize a cross-attention mechanism to develop a method that effectively weights the relationships between diagnoses and treatments. Effectively repre- senting the clinical information contained in medical claims data, this approach generates more representative features for predicting MACE. The ROC-AUC score of our proposed cross-attention-based model was 0.7720, higher than other benchmark models including the conventional atherosclerotic cardiovascular disease model, the light gradient boosting machine, and a self-attention-based model. These results indicate that integrating the clinical structure of medical claims data using a cross-attention mechanism significantly enhances the performance of predictive models.

Wed 9 SeptMachine Learning
The gist
Medical claims have lots of information about patients but it’s often messy and hard for computers to understand well. This study tested a new way to help computer models better learn the connections between patient diagnoses and treatments. By using a method called cross attention, the researchers made predictions about serious heart problems more accurate than previous models. This can help doctors identify patients at risk for major heart events earlier and more reliably.
Open 2609.09824v1

Which medical questions benefit most from detailed answer explanations

Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA

Abstract: Medical question-answering datasets often contain answer labels, whereas high-quality rationales remain scarce, noisy, or costly to validate. This changes the acquisition question: rather than asking which questions should be labeled, we ask which already-labeled questions should receive rationale supervision under a fixed token budget. We study an offline version of this problem in which candidate rationales are visible to the selector but withheld from downstream training unless selected. We propose root-mean-square Robustness-based Sample Prioritization (RMS-RSP), which perturbs hidden states only at rationale tokens and measures the resulting shift in the gold-versus-best-distractor margin. Across five medical QA datasets, MedGemma-4B-IT, three training seeds, ten budgeted non-RSP selectors, and an unbudgeted full-supervision reference, RMS-RSP provides a deliberately qualified result. Its locked-budget accuracy is 60.61% on average versus 60.08% for Random, with a statistically resolved gain only on AfriMed-QA (+1.44 points). Its full-budget accuracy area is not better than Random. However, after three answer-option reorderings, RMS-RSP improves robust accuracy and semantic consistency by 1.91 and 2.85 points on average, respectively, with the same direction on all five datasets. Training on every pool rationale raises macro accuracy to 63.74%, but consumes 29--254 times more rationale tokens and does not uniformly improve robustness. These findings do not establish universal accuracy gains; they instead suggest that rationale-local boundary sensitivity can identify supervision that improves invariance to semantically equivalent formatting changes.

Wed 9 SeptComputation and LanguageArtificial Intelligence
The gist
Medical question-answering systems often label correct answers, but detailed explanations (rationales) are rare and costly to produce. The authors studied how to pick which questions should get these explanations when there's a limited budget. They propose a method that identifies questions where explanations help improve answer consistency, especially when formatting changes. Their results show modest accuracy improvements overall but better robustness to changes in how questions are presented.
Open 2609.09684v1

Clinical diagnosis agents improve safe stopping decisions under risk constraints

Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

Abstract: Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or defer. Existing agent benchmarks largely evaluate accuracy after fixed or unconstrained interaction, leaving autonomous stopping reliability implicit. We present Cros, a risk-constrained stopping layer combining state-wise error ranking, policy design on disjoint development splits, and LTT-style exact tests of selective diagnostic error and minimum autonomous coverage for complete sequential policies. Its finite-sample guarantee requires the candidate family, testing rule, and any randomization to be frozen before calibration labels are accessed. On a 1,834-episode MIMIC-derived abdominal-pain benchmark, the full ranker achieves exploratory state-error AUROC 0.853, compared with 0.715 for maximum class probability and 0.552 for the backbone's native stop score. On the previously viewed 367-episode evaluation split, analytically averaging over the frozen Cros weights yields 16.9% selective error at 78.8% coverage, cost 5.57, and 0.68 tests, versus 30.8% error at 100% coverage, cost 8.14, and 1.53 tests under native stopping. Forced continuation is non-monotone: error is 28.3% with HPI alone and 34.3% after full workup. However, the uniform-weight mixture ablation is cheaper on this viewed split despite missing the locked development margins, and Cros nominally satisfies the joint criterion in only 6 of 20 development resplits. Because evaluation labels were inspected during earlier development, these findings provide exploratory feasibility and audit evidence, not a confirmatory safety certificate.

Wed 9 SeptArtificial Intelligence
The gist
When computer programs help doctors diagnose patients by suggesting tests, they also need to know when to stop asking for more tests and make a diagnosis. The authors developed a new method called Cros to decide when it’s safe to stop testing while keeping errors low. Tested on medical data about abdominal pain, Cros reduced diagnostic mistakes and the number of tests needed compared to earlier methods. This work helps make automated medical diagnosis systems more reliable and safer.
Open 2609.09678v1

Differential privacy improves treatment effect estimates in sensitive data

Differentially Private Average Treatment Effect Estimation by Propensity Score Blocking

Abstract: Average treatment effect (ATE) estimation in observational studies is a fundamental statistical tool used frequently in social science, medicine, and other fields. These fields often work with sensitive data where privacy protections are important, so a differentially private mechanism for ATE estimation is highly desirable. Here we present two propensity score-based algorithms for ATE estimation on observational data, one improving the inverse probability weighting (IPW) method used in prior work, and the other using blocking on the propensity score (BPS). Both show lower error and less bias than prior work, with the BPS-based algorithm frequently reducing error by 75% or more compared to prior work.

Tue 8 SeptCryptography and SecurityMachine Learning
The gist
Estimating how a treatment affects outcomes is important in many fields that handle private data, like medicine and social science. The authors present two new ways to calculate the average effect of a treatment, while protecting individual privacy. Their methods use techniques based on how likely someone is to receive the treatment, called propensity scores. One of their approaches greatly reduces errors compared to previous methods, making estimates safer and more accurate when dealing with sensitive information.
Open 2609.09536v1

Multimodal AI predicts bladder cancer subtypes and progression risks

CHIMERA Challenge Task 2 and 3: Response Subtypes Classification and Progression Survival Prediction in Bladder Cancer Patients using Multimodal Datasets

Abstract: High-risk non-muscle-invasive bladder cancer (HR-NMIBC) carries substantial risks of recurrence and progression, while current clinical risk stratification remains limited. CHIMERA was established as a multimodal AI challenge to benchmark prediction in HR-NMIBC under standardized evaluation. Task BRS predicts RNA-seq-defined BCG Response Subtypes from histopathology and structured clinicopathological data, whereas Task Progression models time-to-progression using histopathology, structured data, and RNA sequencing. A multimodal dataset of 368 patients was divided into public training and hidden validation and test sets. In total, 159 submissions were made, and 13 top-performing models were selected for benchmarking. The best models achieved a weighted F1 score of 0.73 for Task BRS and a C-index of 0.68 for Task Progression. Post-challenge analyses revealed task-dependent modality contributions, cohort-dependent performance degradation, and sensitivity to missing structured data. In Task BRS, histopathology partly compensated for pathology-derived structured variables, whereas progression models showed greater dependence on complementary inputs. Cross-model error analysis further identified patients that were consistently difficult across different architectures, with T1 substage associated with prediction difficulty. These findings highlight barriers to transportability and the importance of missingness-aware modeling and independent multi-institutional validation. CHIMERA provides a standardized multimodal benchmark for bladder cancer and a framework for studying not only model performance, but also robustness, information sufficiency, and patient-level prediction failure.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Bladder cancer can come back or get worse, especially in high-risk cases. The authors set up a challenge called CHIMERA to see how well AI can predict specific cancer types and how quickly the disease might progress using tumor images, patient data, and gene activity. They tested many AI models on a large set of patient data and found that combining different data types helps predictions, though some details still cause errors. Their work shows that it’s important for AI methods to handle missing information and to be tested in different hospitals.
Open 2609.09510v1

Ai model captures full patient health history for better predictions

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

Abstract: The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetime, yet fully exploiting these data to represent and predict patient state trajectories remains a critical challenge. Current AI models often struggle to capture the complex, irregular temporal dynamics and inherent stochasticity of real-world multimodal patient data. Existing AI approaches for modeling longitudinal patient records are predominantly discriminative, limited to a few modalities, constrained by closed categorical vocabularies, treating time as a monotonic inductive bias, or they are limited in forecasting future patient states. We introduce NOAH, a time-aware, task-agnostic, generative transformer model representing and forecasting the full multimodal patient journey. NOAH features a novel bidirectional time integration and a variational latent space to capture the continuous evolution of patient states and the stochasticity of clinical trajectories. Built from over 559 million clinical events from 431,000 hospital visits of 299,000 patients across the MIMIC dataset family, NOAH natively processes medical images, time-series and numeric signals, categorical events, as well as structured and unstructured clinical records. NOAH is the first truly holistic generative model in its field, enabling autoregressive forecasting with optional time control, zero-shot classification, and counterfactual intervention simulation. It generates highly informative and predictive patient state representations that demonstrate strong performance in probing for clinical outcomes, 15 ICD chapters, and 29 comorbidities, as well as in time-to-event prediction. Seamlessly handling diverse modalities and complex temporal dynamics, NOAH provides a versatile, task-agnostic, scalable foundation for intelligent predictive systems in personalized clinical care and digital medicine.

Tue 8 SeptMachine LearningArtificial Intelligence
The gist
Medical records include many kinds of data collected over a patient’s life, but computers have trouble understanding how all this data changes over time. The authors created NOAH, an AI model that looks at medical images, numbers, notes, and events together while paying attention to when things happened. This helps NOAH understand how patients’ health states evolve and make predictions about their future health. It can even simulate 'what-if' scenarios, like what might happen if a patient’s care changes.
Open 2609.09140v1

Personalized models improve disease diagnosis across hospitals nearby

Geographically Regularized AUC-Maximizing Personalized Federated Learning

Abstract: Accurate diagnostic and risk-prediction models are important for supporting clinical decision-making during infectious disease outbreaks. However, privacy and governance requirements may restrict patient-level data sharing across healthcare institutions, and data distributions often vary. Moreover, AUC is widely used to evaluate discriminative performance, motivating its direct optimization in model development. We propose geographically regularized AUC-maximizing personalized federated learning (GrAUC-PFL), which directly optimizes a smooth pairwise AUC surrogate to learn personalized models while keeping patient-level data local and accounting for institutional heterogeneity. Graph-based regularization encourages geographically neighboring institutions to have similar coefficient vectors while retaining a personalized models. Simulations and a real-data application suggest improved discriminative performance, particularly when geographically neighboring institutions have similar data-generating characteristics.

Tue 8 SeptMachine Learning
The gist
During disease outbreaks, hospitals need good computer models to help predict who might be sick. But hospitals usually cannot share detailed patient data with each other because of privacy rules. The authors created a way for hospitals close to each other to build similar but customized models without sharing private data, and these models focus on improving the AUC, a way to measure accuracy. Their method helps predict diseases better, especially when nearby hospitals have similar types of patients.
Open 2609.08379v1

Spfilm improves brain region labeling on pre and post contrast mri scans

Spatial Feature-wise Linear Modulation (SpFiLM) for Contrast Agent-Aware Brain Parcellation

Abstract: Most automated brain parcellation tools are developed and validated on T1-weighted (T1w) MRI. Yet, some clinical workflows for which parcellation is relevant only use contrast-enhanced T1w (T1ce) MRI, on which T1w-trained models are less accurate. We present a unified network that parcellates both pre- and post-contrast agent T1w MRI reliably, trained on a combination of the two with conditioning that spatially modulates its response differently for each. Feature-wise Linear Modulation (FiLM) is a known approach for input-based modulation in networks. It applies a per-channel scale and shift uniformly across the input. However, the appearance change between pre- and post-contrast varies locally across the brain, making FiLM suboptimal for our use case. In this work, we introduce Spatial FiLM (SpFiLM), a conditioning layer whose modulation varies spatially, assembling a voxel-wise scale and shift from image-derived spatial patterns. Using a cohort of 134 patients with paired T1w and T1ce MRI parcellated into 106 classes, the addition of SpFiLM layers in a UNet increased the mean Dice on the test set of 25 patients from 80.2% to 84.1%, a 4.9% relative improvement. Adding SpFiLM layers led to the best performance on both pre- and post-contrast MRI, even when controlling for network parameter counts.

Mon 7 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Different types of MRI scans of the brain can look quite different, especially before and after a contrast dye is used. Models trained on one type may not work well on the other. The authors created a new technique called Spatial FiLM that helps a single model adjust its understanding at each location in the brain depending on the scan type. This approach improved the accuracy of identifying brain regions on both types of scans.
Open 2609.07718v1

Heart transplant prediction models made transparent and easy to audit

Translation of Black-Box Clinical Prediction Models into Standalone Transparent Nomograms: Temporal External Validation in Heart Transplantation

Abstract: We convert black-box clinical prediction models for tabular data into standalone nomograms that can be audited term by term. PRiSM (Partial Responses in Structured Models) takes the shape of each effect and interaction from the source model, not merely which variables mattered, and lets the outcome select and weight them. We tested this in 50,356 heart transplant recipients, with validation in a later era than training. Nomograms from all 5 source models - a public clinical risk score, logistic regression, neural networks, random forests and extreme gradient boosting - met a prespecified noninferiority criterion for discrimination before any further simplification, and generally preserved calibration and clinical net benefit. Those from the 3 machine-learning models showed no detectable difference in discrimination from de novo generalized additive and explainable boosting models, exceeded neural additive models, and carried fewer terms than the explainable boosting model. PRiSM is released as an open-source Python package.

Mon 7 SeptMachine Learning
The gist
Many clinical prediction models are like 'black boxes'—they give results but don’t explain how. The authors created a method called PRiSM that turns these complex models into clear, straightforward charts called nomograms. These nomograms show how each factor affects the prediction, making them easier to understand and check. They tested this system on heart transplant data and found the transparent models worked just as well as the original ones. They also shared the method as an open-source tool anyone can use.
Open 2609.07610v1

Large language models struggle to use long patient records for clinical decisions

ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making

Abstract: The application of large language models (LLMs) to personalized medical assistants has garnered growing interest. However, existing medical benchmarks largely rely on static question answering with pre-selected evidence, leaving unclear whether LLMs can make reliable clinical decisions from real longitudinal electronic health records (EHRs). To bridge this gap, we introduce ObGynLongBench, a rule-grounded long-context EHR benchmark for obstetric and gynecologic decision-making, comprising 1,500 clinical decision-point cases from 976 real pregnancy EHR histories and traceable rules. Each case is anchored to a patient, a pregnancy-timeline point, and a pre-decision information boundary, enabling Evidence-only, Visit-level EHR, and History-level EHR evaluation. Evaluating 17 LLMs reveals a substantial Evidence-to-EHR Gap: models perform well when evidence is directly provided, but accuracy drops when evidence must be extracted from same-day records or full pre-decision EHR histories. Further analyses identify evidence utilization as a key bottleneck: performance decreases with longer EHR contexts and more complex evidence requirements, and earlier failures often predict later failures within the same patient history. Finally, active-search agents perform best among EHR access strategies, highlighting patient-specific evidence utilization as a central challenge for reliable personalized medical assistants. Resources are available at https://github.com/xiangjun2003/ObgynLongbench.

Mon 7 SeptComputation and LanguageArtificial Intelligence
The gist
Large language models have shown promise in answering medical questions when given the right facts. However, this paper shows that when these models need to make decisions based on a patient’s full medical history over time, especially for pregnancy care, they perform much worse. The authors created a new test with real pregnancy records to measure this and found that models often fail to find and use the right evidence from long health records. This suggests more work is needed to build reliable medical assistants for ongoing patient care.
Open 2609.07601v1

Machine learning benchmark compares models on biomedical tables

TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables

Abstract: Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. We introduce TabBench-Bio, a living and interactive benchmark of 43 biomedical datasets spanning multiple domains. Under a shared cross-validation protocol, we compare classical estimators, neural networks, and tabular foundation models across 28 feature-by-sample operating points. At the reference cell of 10,000 features and 100 training samples, RealTabPFN v2.5 has the highest point estimate, followed by Logistic Regression and TabDPT, whose point estimates are nearly identical. A paired bootstrap over the target pool separates RealTabPFN v2.5 from Logistic Regression by 145 Elo (95% interval [59, 232]). Tabular foundation models generally occupy the leading ranks, while the strongest configuration depends on the operating point and biomedical modality. The AutoML framework AutoGluon, using its one-hour "extreme" preset, is configured as a separate resource-intensive reference and is reported here at the reference cell. Fold-level predictions, run status, and deterministic aggregations make every reported result reproducible and reusable. We invite the community to contribute: TabBench-Bio is designed to grow, and we welcome submissions of new biomedical tabular datasets, particularly from underrepresented assays and clinical endpoints, for inclusion in future releases. The interactive leaderboard is available at: https://tabbench-bio.eu

Mon 7 SeptMachine LearningArtificial Intelligence
The gist
Biomedical data tables often have many features but only a few samples, making it hard to test machine learning methods fairly. The authors created TabBench-Bio, a growing collection of 43 biomedical datasets to compare different machine learning methods under consistent conditions. They found that specialized tabular models generally perform best, but the top method depends on the data type and sizes. This benchmark is interactive and invites others to add datasets to better cover underrepresented biomedical areas.
Open 2609.07441v1

Human feedback improves AI safety in clinical out-of-distribution detection

CHILD: Human-in-the-Loop OOD Detection for Safe Clinical Deployment

Abstract: Out-of-distribution (OOD) detection is critical for safe deployment of medical AI systems. Recently, test-time adaptation (TTA) has emerged as a new paradigm for OOD detection, automatically adjusting detector behavior during deployment. However, such automatic adaptation mechanisms may raise safety concerns in safety-critical clinical environments. While physician oversight can mitigate these risks, it is resource-intensive and must be judiciously allocated. To reconcile safety with efficiency, we propose CHILD, a training-free framework designed to enhance streaming OOD detection via sparse human feedback. Operating under strict budget constraints, CHILD employs an adaptive risk-aware sample selection mechanism to pinpoint only the most decision-uncertain samples for review. Crucially, it maximizes the utility of this sparse feedback through a retrieval-based score calibration module, which refines model predictions using a compact feature cache without any parameter updates. Extensive experiments on four medical benchmarks demonstrate that CHILD turns limited supervision into significant reliability gains: with a sparse feedback budget of only 5%, it reduces the average FPR95 from 72.63% to 60.26% and improves AUROC from 75.53% to 81.85%, consistently outperforming state-of-the-art baselines. Our code is publicly available at https://github.com/figec/CHILD.

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
AI systems in medicine can make mistakes when they see unusual cases different from what they were trained on. The authors developed CHILD, a system that asks doctors to check only the most uncertain cases during use, saving time. CHILD then uses those few expert inputs to improve its predictions without needing complex retraining. This approach helps make AI tools more reliable and safer for clinical use.
Open 2609.07188v1

LLM framework extracts lung cancer tumor stages with transparency

SIFTING: A Novel LLM-Based Framework for Structured and Transparent Information Extraction from Clinical Free-Text Reports, with Application to Tumor Staging in Lung Cancer

Abstract: Background: Large language models (LLMs) show promise for extracting information from clinical free-text documents, but their outputs are often unstructured and lack traceability, complicating validation and adoption in clinical workflows. In this work we introduce SIFTING, an LLM-based framework designed to address these shortcomings. Methods: SIFTING combines the language comprehension capabilities of LLMs with segment-level processing and structured prompts with strict output control, linking findings to the source text to enable both accurate and transparent information extraction. To demonstrate its capabilities, we applied the framework to the task of extracting tumor T-stage information from 130 lung cancer radiology reports (SIFTING-T-stage). A compact 4-bit quantized version of the open-source LLM Llama-3.3-70B (35 GB) was used in a fully self-hosted setup, providing full control over data and model. Performance was evaluated against a reference standard created by four clinical experts and compared with a range of LLMs as used in a conventional single-prompt approach, using bootstrap resampling to estimate confidence intervals. Results: SIFTING-T-stage achieved an accuracy of 90% (95% CI: 84-95) against the reference standard. We found its performance to be comparable to even the largest state-of-the-art LLMs with reasoning capabilities and to be interchangeable with clinical experts (p < 0.001), while at the same time offering full traceability through source text references. Conclusion: SIFTING enables accurate, structured, and traceable information extraction from clinical free-text documents. It ensures data control, reproducibility, and verifiable outputs that can support clinical validation and workflow integration.

Mon 7 SeptComputation and Language
The gist
Extracting detailed medical information from free-text clinical reports is hard because such data is often unstructured and difficult to verify. The authors created a system called SIFTING that uses advanced language models to carefully pull out specific tumor stage details from lung cancer reports, while linking each fact back to where it appears in the text. This makes the output more accurate and easier to check. Their system performed as well as experienced clinicians at this task, using a self-hosted language model that preserves data control.
Open 2609.07185v1

Benchmark tests show challenges in reasoning over privacy-protected records

CIPHER: Benchmarking Cross-record Inference over Privacy-Hardened Evidence Records

Abstract: Reasoning over privacy-constrained records requires combining structured attributes with evidence from free-text narratives. We introduce CIPHER (Cross-record Inference over Privacy-Hardened Evidence Records), a benchmark of expert-validated questions from consumer-finance, clinical, and law-enforcement records. The questions cover common tabular operations and include executable SQL supervision. We evaluate retrieval, prompting, table-specialist, and hybrid symbolic-neural systems under native redaction and surrogate-based evidence restoration. All system families exhibit substantial failures even when supporting records are provided. Most errors arise from incorrect record selection and predicate interpretation rather than arithmetic execution. Privacy transformations have non-uniform effects, sometimes obscuring necessary evidence and sometimes reducing distraction. CIPHER provides a reproducible testbed for diagnosing these failures and assessing how transformations of sensitive text affect reasoning over hybrid records.

Mon 7 SeptCryptography and SecurityArtificial Intelligence
The gist
Dealing with private records means some details are hidden or changed, making it hard to answer questions that mix tables and free-text information. This paper presents CIPHER, a set of expert-made questions from finance, health, and law records to test how well systems handle this problem. The authors found that many systems struggle, especially with choosing the right information and understanding conditions, even when given the right records. They also noticed that privacy protections can both hide needed facts and remove distractions, affecting performance in different ways. CIPHER helps others study these challenges and improve reasoning over sensitive, mixed-format records.
Open 2609.07022v1