Why Large Language Models Fail at Tabular Prediction
2026-08-03 • Machine Learning
Machine Learning
AI summaryⓘ
The authors investigate why large language models (LLMs), which are great at many tasks, struggle with predicting data from tables. They test five ideas for this failure and find that the main problem is how LLM accuracy drops as the number of data features (dimensions) grows, unlike traditional methods that handle this well. In low dimensions, LLMs act like local, distance-based predictors, but their behavior in higher dimensions is unique and not matched by usual models. The authors highlight that while LLMs falter on tabular data, the exact way they make predictions remains unclear.
Large Language ModelsTabular DataPredictive AnalyticsDimensionalityLinear ProjectionsTokenizationLocal ModelsClassical Machine LearningCSV Format
Authors
Marta Garnelo, Wojciech M. Czarnecki
Abstract
Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. This gap is the founding premise of the fast-growing field of tabular foundation models, but the question of why generic LLMs fail has remained open. We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning - and systematically evaluate five hypotheses for the failure: (a) an inability to handle noisy or non-linearly-separable data; (b) the linearised CSV format obscuring column structure; (c) the tokenisation of numeric values; (d) the number of test points classified per query; and (e) the dimensionality of the input. Controlled experiments falsify (a)-(d). Dimensionality, in contrast, is decisive: sweeping random linear projections of thirty-one benchmark datasets, the LLM is the only method among nine whose accuracy decreases as dimensionality grows, while every classical baseline stays flat or improves. A behavioural comparison against 252 configured classical models finds that in two dimensions the LLM predicts like a local, distance-based method (up to 91.6% grid agreement), but in higher dimensions no classical model - even when augmented with tuned, dimension-dependent noise - reproduces its predictions. We do not claim to have identified the internal mechanism; our results show, more modestly, that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.