Papers for

biomedical data engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

AI agents struggle to reliably analyze data and generate novel scientific hypotheses

DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

Abstract: Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: https://huggingface.co/datasets/discern-bench-anon/discern-benchmark

Sun 27 SeptArtificial Intelligence
The gist
Automated AI systems that do scientific research need to check if data is trustworthy, verify their analysis, and create new ideas based on real evidence. The authors created a benchmark called DISCERN to test all these steps together using real public datasets. They found that while AI can do some parts well, it often rejects good data, misses problem details when drawing conclusions, and produces inconsistent new hypotheses. Therefore, current AI agents are not yet able to perform scientific discovery on their own reliably.
Open → 2609.33357v1

Federated learning improves missing clinical data filling across centers

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Abstract: Multi-center clinical studies and biomedical research collaborations increasingly seek to utilize data across centers to build models that generalize beyond any single center. This creates two distinct challenges: data protection regulations may restrict the sharing of raw patient data across institutions, while centers may collect only partially overlapping sets of features under different protocols. Federated learning enables collaborative model training without centralizing raw data. However, existing federated imputation methods rarely evaluate feature-level missingness, in which entire features are unobserved at some centers. To address this setting, we adapt the ReMasker masked autoencoder to federated learning (Fed-ReMasker), enabling centers to impute features never observed locally by leveraging knowledge learned across collaborating centers. We evaluate Fed-ReMasker in a benchmark spanning synthetic datasets with linear and nonlinear relationships and real-world tabular datasets, including clinical data. The benchmark varies the number of centers, the missingness ratios, and client heterogeneity. Fed-ReMasker achieves the lowest imputation error in 93.2% of value-level and 96.7% of feature-level scenarios in the homogeneous benchmark. It also remains robust to client heterogeneity using simple federated averaging, outperforming all baselines in all 36 value-level scenarios and each baseline in at least 35 of 36 feature-level scenarios, and comes within 3.0% on average of a centralized model trained on the pooled data.

Wed 23 SeptMachine LearningArtificial Intelligence
The gist
When hospitals and research centers collect different types of patient data but want to use it together, sharing raw data is often not allowed and some data features may be entirely missing in some places. The authors created Fed-ReMasker, a tool that helps predict and fill in missing data without sharing sensitive patient information, by learning patterns from all involved centers. Their tests show that this method guesses missing values more accurately than other existing methods, even when some centers have very different data. It works nearly as well as if all data were combined in one place, which isn’t usually possible.
Open → 2609.28105v1

Biomedical entity linking improves with detailed token matching

BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval

Abstract: Biomedical Entity Linking disambiguates mentions to entities in a knowledge base (KB), making it the cornerstone of information extraction pipelines. While embedding-based models are a popular approach for the task, they suffer from a key limitation. They compress mentions (and entities) into a single vector, forcing the model to average away crucial fine-grained differences. We present BELXTR, a novel embedding model based on the multi-vector (a.k.a. late interaction) architecture, which allows to leverage token-level matching information. BELXTR extends the original XTR model to biomedical entity linking by integrating an existing task-specific training objective and exploring active query expansion. Experiments across ten corpora and five KBs show that BELXTR improves upon current state-of-the-art in half of the corpora with an average improvement of 5pp recall@1. The largest gains are reported on the challenging cross-species gene disambiguation subtask, where BELXTR outperforms an LLM-powered retrieve-and-rerank pipeline and closely approaches a specialized rule-based system. Our results highlight multi-vector models as a practical alternative to hard-to-maintain rule-based systems or in scenarios where LLM-based reranking is too costly as in PubMed-scale mining. The code to reproduce our experiments can be found at: https://github.com/sg-wbi/belxtr.

Tue 22 SeptComputation and Language
The gist
Biomedical entity linking helps computers figure out exactly which medical terms in text refer to which specific entries in a biomedical knowledge database. The authors found that many current methods oversimplify this task by turning each term into a single summary, losing important details. They created a new method called BELXTR that looks at individual word parts within the terms, improving accuracy. Their tests showed it works better on many datasets, especially when distinguishing similar genes across species.
Open → 2609.25859v1

Single-cell models reveal challenges with rare cell type recognition

Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions

Abstract: Single-cell foundation models (scGPT, scBERT, Geneformer) achieve cell-type classification accuracy up to 97.5% in our experiments, yet this aggregate accuracy can mask systematic failure on rare, often disease-relevant cell populations that long-tail loss functions are widely assumed to address. We present a systematic benchmark of six long-tail loss functions (cross-entropy, weighted CE, class-balanced loss, focal loss, LDAM, logit-adjusted softmax) across three architectures and three datasets (Multiple Sclerosis, Zheng68K, human Pancreas), totaling 162 controlled training runs (3 backbones x 3 datasets x 6 losses x 3 seeds). The gap between overall accuracy, Macro-F1, and rare-class recall under plain cross-entropy is consistent across all nine (architecture, dataset) settings, driven by dataset structure rather than pretraining. Rare-class failure itself splits into two regimes with distinct embedding-geometry signatures, visible before any loss is chosen: some classes are recoverable by the right loss, while others retain linear separability yet are absorbed into unrelated classes' neighborhoods under every evaluated loss and architecture. Among the recoverable classes, the efficacy of reweighting is predicted by a class's absolute training-set size, rather than its share of the dataset or the dataset's overall imbalance ratio. Class-balanced loss and LDAM are the most consistent choices across all nine settings, while logit adjustment trades rare-class precision for recall rather than improving both. Our results give both a reusable benchmark and mechanism-grounded practical guidelines for combining foundation models with imbalanced biological data.

Sun 20 SeptMachine Learning
The gist
Classifying cell types from biological data is easier for common cells but tougher for rare ones, which may be important in disease. The authors tested different methods designed to help detect these rare cells across various models and datasets. They found that some rare cell types can be identified better with certain loss functions, but others remain hard to separate correctly. Also, the usefulness of techniques that give more weight to rare classes depends on the actual number of rare training examples, not just their proportion. This work helps guide better use of AI models with imbalanced biological data.
Open → 2609.23325v1

Scientific knowledge graphs improve distributed vector search efficiency

COMPASS: Steering Distributed Vector Search with Scientific Knowledge Graphs

Abstract: Vector databases use hashing to partition data across "shards," logical units for distributed execution. This placement, however, destroys semantic locality, forcing each query into scatter-gather limited by the slowest shard. Vector-space clustering can help, but scientific evidence is often connected by factual relations that do not align with embedding distance. We present COMPASS, a framework that uses a knowledge graph (KG) to determine data placement and query-time shard selection. COMPASS detects communities, splits oversized communities, inserts embeddings by subject entity, and routes queries to a small set of shards. Across four biomedical KGs, our method searches only 13-18% of the corpus while preserving broadcast recall and recovering up to 2.6x more multi-hop evidence than an embedding-based baseline. On 15 HPC nodes, COMPASS sustains 7.9x higher throughput with lower tail latency than hash-based broadcast. These results show that KG structure provides a compact complement to embedding geometry for scalable vector search.

Fri 11 SeptDatabasesDistributed, Parallel, and Cluster Computing
The gist
Searching large collections of data with vector databases can be slow because data is split randomly across many parts, causing queries to check everything. The authors show that using scientific knowledge graphs, which capture relationships between facts, can better group data and guide queries to fewer relevant parts. Their approach, COMPASS, checks less data while still finding the important connections between pieces of information. This leads to much faster searching without losing important results.
Open → 2609.13452v1