Availability of unnecessary tools reduces large language model answer accuracy

When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering

Computation and Language

Summary

Sometimes big language models are given helpful tools to find answers, but this study shows that if an extra tool is available that isn’t needed, the model’s ability to answer questions from its own knowledge gets worse. This effect happens even when the tool is hardly ever used. The researchers tested six language models across many questions and found accuracy dropped a lot when unnecessary tools were present. However, a simple instruction telling the model when to use tools helped fix most of this issue.

What this means in practice

  • For ai system developers: Improve language model answer accuracy by adding scope-aware instructions that prevent unnecessary tool use during question answering.
  • For chatbot platform engineers: Design chatbot interfaces that carefully manage tool availability to avoid degrading response quality on questions answerable from the model’s knowledge alone.

Authors

Saanvi Paturi, Arsen Kenzhebayev, Arham Sethi, Vyas Raina, Ivaxi Sheth, Vatsal Raina

Abstract

Large language models (LLMs) are increasingly deployed with external tools that extend what they can do beyond their own knowledge. Tools help on tasks that need external information, but their availability may also change how a model handles questions that do not need them. Prior work has mostly asked whether models select and use tools appropriately; whether an unnecessary tool changes the correctness of answers has received less attention. We ask whether making a related but unnecessary tool available affects a model's ability to answer from its own knowledge, and whether a preceding tool interaction changes this behaviour. We construct 500 query pairs across 10 knowledge domains. Each pair consists of a tool query, which needs the domain's tool, and a closed-domain query, which does not. Six LLMs are evaluated with the tool unavailable, available, and available after a prior tool call. Across 3,000 baseline trials the pooled answer rate is 98.2%. When an unnecessary tool is available it falls to 63.5%, with large differences between models. The decrease occurs even when the tool is rarely called, so it cannot be explained by unnecessary tool invocation alone. A one-sentence scope-aware system instruction recovers most of the lost answers.