Text to SQL model accuracy improves using plans across database dialects

Closing the Cross-Dialect Gap: Query Plans as a Portable Interface in Text-to-SQL

Computation and Language

Summary

Text-to-SQL models usually learn one type of database language, which causes errors when used with others. The authors show this is a big problem for many models. They suggest having the model produce a general query plan, which can be turned into any specific database language afterwards. This method improves accuracy across many SQL languages without hurting performance on the original language, and sometimes even improves it. They also created a new way to fairly measure correctness across different database languages.

What this means in practice

  • For database system developers: Enable text-to-SQL tools to generate queries that work across multiple database dialects by using dialect-agnostic query plans.
  • For enterprise software engineers: Improve the reliability of natural language database querying in products that support varied SQL databases by integrating plan-based translation.

Authors

Corentin Royer, Robin Oester, Yotam Perlitz, Yannick Metz, Andrea Giovannini, Mennatallah El-Assady

Abstract

Text-to-SQL systems are typically trained and evaluated on a single dialect (SQLite), yet production deployments span PostgreSQL, MySQL, ClickHouse, and beyond. We show that this single-dialect assumption leads to a substantial drop in cross-dialect accuracy for every model we tested. The drop persists across scale, architecture, and even purpose-built text-to-SQL systems. We argue that the fix is to change the generation target: instead of asking an LLM to emit dialect-specific SQL, we have it emit a dialect-agnostic relational algebra query plan, which a deterministic compiler then renders into SQL for any supported backend. Across thirteen models from 3B to frontier scale, this restores cross-dialect portability nearly uniformly, at a small cost in peak accuracy on the model's home dialect for capable prompted models and none once fine-tuned on plans; under matched fine-tuning, plan supervision yields a stronger model than SQL supervision. We also introduce MetricName, a question-aware result-set comparator needed to evaluate fairly across dialects, where existing metrics confound semantic errors with benign cross-dialect variation. More broadly, the result is a reminder that a generation target chosen for execution is not necessarily the one that maximizes generation quality.