• Sources: preprint, discussion
  • Summary: Marta Garnelo and Wojciech M. Czarnecki, submitted 2026-08-03, study a frontier LLM in a single generation pass over a prompt holding the full training and test data, with no tools, no agentic scaffolding and no fine-tuning. They test five hypotheses for the known failure on tabular prediction and report that controlled experiments falsify four of them: noisy or non-linearly-separable data, the linearised CSV format obscuring column structure, numeric tokenization, and the number of test points per query. Sweeping random linear projections of 31 benchmark datasets, they report the LLM is the only method among nine whose accuracy decreases as dimensionality grows, while every classical baseline stays flat or improves. Compared against 252 configured classical models, they report the LLM predicts like a local distance-based method in two dimensions with up to 91.6% grid agreement, and that in higher dimensions no classical model reproduces its predictions even with tuned dimension-dependent noise. The authors explicitly do not claim to have identified the internal mechanism. This is a preprint with no independent reproduction.
  • Why it matters: Prompt formatting and numeric encoding are not the lever, so the choice between an LLM and a classical model on a wide table turns on column count.

send feedback on this story