The Bitter Lesson of Tool Calling

2026-08-06Computation and Language

Computation and Language
AI summary

The authors studied how different ways of letting language models use tools affect their performance. They compared a newer method called programmatic tool calling, where models write and run Python code to use tools, to the older method that uses fixed JSON commands. Testing 14 models, they found that programmatic tool calling usually worked better or at least as well as JSON commands, especially with newer models like GPT-5.6. This method also stayed reliable even in harder situations where the old method struggled. Overall, the authors show that programmatic tool calling is a strong and flexible way for models to interact with tools.

Large Language ModelsProgrammatic Tool CallingJSON Tool CallingPython StubsBFCL BenchmarkParallel Fan-outContext RotGPT-5.6Agent-based ModelsTool Use in AI
Authors
Ishan Patel, Sahil Sen, Elias Lumer, Vamse Kumar Subbiah
Abstract
Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by replacing rigid JSON calls with scripts that chain and parallelize naturally. However, a systematic evaluation of tools as code on an established benchmark across current and prior model generations under real-world task conditions has not been conducted. In this work, we empirically compare programmatic tool calling (PTC) to native JSON tool calling across 14 language models on BFCL v4. In the programmatic tool calling paradigm, tools are exposed as typed Python stubs that the model invokes through code, with execution and results handled in a single agent turn. Programmatic tool calling matches or exceeds native JSON tool calling in 11 of 14 models on BFCL v4, with the GPT-5.6 family achieving a 10.6% improvement over the JSON tool calling baseline. Further, it matches or outperforms baseline in 13 of 14 models under parallel fan-out, and holds stable under context rot conditions where baseline degrades 2.3% on average. Our results demonstrate that programmatic tool calling is a viable and robust alternative to JSON tool calling, with performance tracking model capability across release generations.