Biomedical entity linking improves with detailed token matching
BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval
Computation and Language
Summary
Biomedical entity linking helps computers figure out exactly which medical terms in text refer to which specific entries in a biomedical knowledge database. The authors found that many current methods oversimplify this task by turning each term into a single summary, losing important details. They created a new method called BELXTR that looks at individual word parts within the terms, improving accuracy. Their tests showed it works better on many datasets, especially when distinguishing similar genes across species.
What this means in practice
- •For biomedical data engineers: Improve automatic linking of biomedical terms to curated databases in text mining pipelines at scale.
- •For pharmaceutical intelligence analysts: Enhance identification of gene and protein mentions across species in literature without relying on costly large language models.
Authors
Samuele Garda, Ulf Leser
Abstract
Biomedical Entity Linking disambiguates mentions to entities in a knowledge base (KB), making it the cornerstone of information extraction pipelines. While embedding-based models are a popular approach for the task, they suffer from a key limitation. They compress mentions (and entities) into a single vector, forcing the model to average away crucial fine-grained differences. We present BELXTR, a novel embedding model based on the multi-vector (a.k.a. late interaction) architecture, which allows to leverage token-level matching information. BELXTR extends the original XTR model to biomedical entity linking by integrating an existing task-specific training objective and exploring active query expansion. Experiments across ten corpora and five KBs show that BELXTR improves upon current state-of-the-art in half of the corpora with an average improvement of 5pp recall@1. The largest gains are reported on the challenging cross-species gene disambiguation subtask, where BELXTR outperforms an LLM-powered retrieve-and-rerank pipeline and closely approaches a specialized rule-based system. Our results highlight multi-vector models as a practical alternative to hard-to-maintain rule-based systems or in scenarios where LLM-based reranking is too costly as in PubMed-scale mining. The code to reproduce our experiments can be found at: https://github.com/sg-wbi/belxtr.