Papers for

enterprise software engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Text to SQL model accuracy improves using plans across database dialects

Closing the Cross-Dialect Gap: Query Plans as a Portable Interface in Text-to-SQL

Abstract: Text-to-SQL systems are typically trained and evaluated on a single dialect (SQLite), yet production deployments span PostgreSQL, MySQL, ClickHouse, and beyond. We show that this single-dialect assumption leads to a substantial drop in cross-dialect accuracy for every model we tested. The drop persists across scale, architecture, and even purpose-built text-to-SQL systems. We argue that the fix is to change the generation target: instead of asking an LLM to emit dialect-specific SQL, we have it emit a dialect-agnostic relational algebra query plan, which a deterministic compiler then renders into SQL for any supported backend. Across thirteen models from 3B to frontier scale, this restores cross-dialect portability nearly uniformly, at a small cost in peak accuracy on the model's home dialect for capable prompted models and none once fine-tuned on plans; under matched fine-tuning, plan supervision yields a stronger model than SQL supervision. We also introduce MetricName, a question-aware result-set comparator needed to evaluate fairly across dialects, where existing metrics confound semantic errors with benign cross-dialect variation. More broadly, the result is a reminder that a generation target chosen for execution is not necessarily the one that maximizes generation quality.

Sun 27 SeptComputation and Language
The gist
Text-to-SQL models usually learn one type of database language, which causes errors when used with others. The authors show this is a big problem for many models. They suggest having the model produce a general query plan, which can be turned into any specific database language afterwards. This method improves accuracy across many SQL languages without hurting performance on the original language, and sometimes even improves it. They also created a new way to fairly measure correctness across different database languages.
Open → 2609.33670v1

Page images improve document QA accuracy but increase latency and cost

Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA

Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, judge, and scoring pipeline fixed, across four commercial model endpoints, two corpora, and two context regimes (gold evidence pages and the full document). On documents that fit the image budget, page images lead on accuracy at every document length on both corpora, but this advantage carries a growing latency and cost premium: text latency stays roughly flat as documents lengthen while image latency rises steadily. Text and images also fail on different questions, with exactly one representation correct on 19--25% of items across the reported cells, so neither subsumes the other. Exploiting this complementarity, a lightweight TF-IDF router that reads only the question text gains 2.6 points over always-text while cutting median latency 30% relative to always-vision, on a document-disjoint held-out split.

Fri 18 SeptArtificial Intelligence
The gist
When computers answer questions about documents, they can use either the words extracted from the pages, images of the pages themselves, or both. This paper shows that using page images generally helps computers answer more accurately but takes more time and computing power as documents get longer. Interestingly, some questions are answered correctly only by looking at images or only by looking at text, so using both can be better. The authors also built a simple system that chooses between text or image input based on the question, improving speed and accuracy.
Open → 2609.22628v1

Small models coordinate to solve complex tasks without cloud reliance

OrchSLM: Probing the Dynamics of Small Language Model Orchestration

Abstract: Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. Small language models (SLMs) offer a compelling alternative: recent studies suggest that many repetitive and narrowly scoped subtasks in agentic workloads may be better served by specialized SLMs than by monolithic LLMs. However, the limited capacity and context windows of SLMs can constrain long-horizon reasoning and interaction-heavy orchestration strategies such as iterative verification and debate. This motivates a complementary, non-interactive paradigm in which heterogeneous SLMs independently generate candidate solutions and a router orchestrates their cached samples without further model interaction. To further understand the mechanisms of such orchestration, we introduce OrchSLM, a routing framework that unifies existing non-interactive orchestration methods and exposes their underlying design choices as controllable parameters. Using OrchSLM as a systematic probe, we reveal how orchestration behavior emerges from diverse knobs, including the task structure, model-pool composition, and multi-agent consensus.

Fri 11 SeptArtificial Intelligence
The gist
Using big language models can be slow, expensive, and require constant internet. The authors explore how smaller language models can each work on parts of a problem independently and then get combined by a system that picks the best answers without the models talking to each other. They introduce OrchSLM, a way to study how different settings affect this coordinating system. This helps us understand how smaller models can team up effectively without needing lots of computing power or online connection.
Open → 2609.13470v1

Graph rAG improves code migration by preserving structure and dependencies

Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration

Abstract: As enterprises modernize legacy monolithic systems to microservices, Large Language Models (LLMs) are heavily utilized for automated code translation. However, traditional vector-based Retrieval-Augmented Generation (Standard RAG) struggles to capture topological relationships. It fetches isolated chunks that sever inheritance chains, leading to high compilation failure rates. This paper introduces a Hierarchical Context-Resident Graph (HCRG) methodology to resolve these limitations. Our pipeline uses tree-sitter for Abstract Syntax Tree (AST) extraction, maps architectural edges into a Google Cloud Spanner Property Graph, and serializes this structure into a Gemini Context Cache for topological, parent-first code translation. We shift evaluation from naive text-overlap to a custom 7-metric Software Engineering framework. Traditional metrics like CodeBLEU (which scored 91% for both methods) effectively masked Standard RAG's structural failures behind syntactically plausible but broken code. Empirically, Graph RAG decisively mitigates dependency loss: API hallucination rates dropped from 56.4% to 16.2%, Dependency Resolution Quality improved from 34.8% to 65.9%, and Parent-Child Consistency rose from 26.7% to 45.5%. However, Graph RAG introduces specific trade-offs. The dense global context causes defensive over-engineering by the LLM, reducing Cyclomatic Complexity Consistency from 71.6% to 46.7%, and slightly degrades Docstring Preservation (67.0% to 61.0%). Ultimately, while trading code complexity for reduced hallucinations, Graph RAG provides a substantially more viable, architecturally sound path for automated enterprise codebase modernization.

Fri 11 SeptArtificial Intelligence
The gist
Modernizing old software into smaller, manageable parts is hard, especially when automatically translating code with AI. The authors show that the usual method which treats code pieces separately misses important connections, causing many errors. They introduce a new approach that understands the code's structure better by using graphs to keep track of these connections. This new method makes the translated code more reliable, although sometimes it creates slightly more complex code than before.
Open → 2609.12464v1

Retrieval-confidence layer detects missing context in enterprise code generation

RCL: A Retrieval-Confidence Layer for Detecting Insufficient Context in Enterprise Retrieval-Augmented Code Generation

Abstract: Retrieval-Augmented Generation (RAG) for code generation has been studied extensively on public repositories, where a model's parametric knowledge often compensates for imperfect retrieval. This breaks down in enterprise codebases, where private APIs, internal frameworks, and undocumented team conventions fall entirely outside any model's pretraining distribution. Recent work on private-library code generation shows that even oracle (perfect) retrieval does not eliminate errors, but locates failures downstream in API usage; separately, confidence-gated retrieval has been studied for open-domain question answering using model-internal confidence. Neither addresses whether retrieval itself was structurally sufficient for a private-code query before generation begins. We introduce RCL (Retrieval-Confidence Layer), a lightweight module inserted between retrieval and generation that combines a call-graph-derived structural coverage score with a novelty score estimating a query's dependence on knowledge outside the model's prior, to detect insufficient retrieval before generation occurs. When confidence falls below a calibrated threshold, RCL triggers a targeted follow-up retrieval or labels the output for human review, rather than generating silently against incomplete context. We describe RCL's architecture, formalize its scoring functions, and propose an evaluation methodology using a private-code benchmark built by injecting synthetic internal APIs into open-source Java repositories, simulating the enterprise condition without proprietary code. We report results (Section 7) comparing RCL against similarity-only retrieval on generation correctness. Our position is that retrieval sufficiency, assessed structurally rather than from model-internal confidence, is a distinct and currently underaddressed signal for building safer code-generation systems in private, enterprise settings.

Thu 10 SeptSoftware Engineering
The gist
Generating code with AI often relies on retrieving related code snippets, but in private company codebases, this retrieval misses important internal or undocumented parts. The authors propose a new method called RCL that checks whether the retrieved code covers enough relevant parts before the AI tries to write new code. If it detects missing pieces, it can call for more searching or ask for human help instead of risking wrong code. They tested RCL using a simulation with injected internal APIs to mimic private code environments.
Open → 2609.11023v1