Large language models show big differences using similar tool interfaces
Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias
Artificial Intelligence
Summary
When large language models (LLMs) use tools through different but equivalent interfaces, they often perform inconsistently. This problem, called schema bias, means that changing how a tool’s actions are presented can cause the same model to fail or succeed dramatically. The authors tested many LLMs with variations of tool interfaces that do the same thing but look different, finding large performance swings. They also showed that training a model helps only if the exact altered interface is in the training data. This reveals that LLMs do not fully understand the underlying tasks but rely heavily on how tool actions are shown.
What this means in practice
- •For ai system developers: Design more robust LLM agents by testing tool interfaces beyond a single schema to avoid unexpected performance drops.
- •For software engineers building ai assistants: Use sample query evaluations to identify difficult tool schemas before deployment, improving reliability of AI assistant tool integration.
Authors
Yinhong Liu, Zhili Tan, Zilin Wang, Zhijiang Guo
Abstract
Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool definitions, and an agent that has truly learned a task should behave consistently across them. We show that current agents often do not, a phenomenon we term schema bias. To study this systematically, we introduce an executable transformation framework that rewrites a native tool schema using nine operators, including merging and splitting tools, altering how a single tool is expressed, and distributing one action across several dependent calls. The tasks, executable actions, and reachable states remain fixed, so any change in success is attributable to the interface alone. Evaluating eleven LLMs, including two closed models, on up to 32 schema variants, we ask how large schema bias is, how it manifests, whether the difficulty of a schema variant can be predicted without a full evaluation, and whether training removes it. We find that schema bias is substantial even for the newest models: success rates range from complete failure to 97% depending solely on the schema. To reliably estimate schema difficulty, it requires running a small sample of the target queries. Training repairs a schema variant only when that variant appears in the training data.