WildSEEK: Evaluating Language Models for Information-Seeking
2026-08-31 • Computation and Language
Computation and LanguageComputers and Society
AI summaryⓘ
The authors created WildSEEK, a dataset of 3,000 real user questions about topics like health and finance, to better test how language models answer information requests. They also made tools to evaluate these answers, especially for complex questions requiring analysis rather than just facts. By studying over 1.8 million queries, they found many questions involve risks and that language models often fail by agreeing too much, relying on limited perspectives, or mishandling sensitive groups. Their work helps measure the safety, fairness, and reliability of language models when they provide information to people.
language modelsinformation-seeking queriesrisk-sensitive domainsfactoid queriesanalytical queriesevaluation frameworksycophantic behaviorUS-centric biasvulnerable populationsquery annotation
Authors
Tanise Ceron, Joachim Baumann, Elisa Bassignana, Berat Cabuk, Dirk Hovy, Debora Nozza
Abstract
Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of "in the wild" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our findings show that LLM responses fail more often in four criteria: sycophantic behavior, overreliance, a default US-centric perspective, and poor handling of vulnerable populations -- with failure rates being mostly higher for analytical queries. By providing methods to monitor the reliability, safety, and fairness of LLM behavior, our dataset and evaluation framework offer an empirical foundation for the broader question of how these systems should behave as they take on a growing role in information access.