A large language model is, mechanically, a predictor of the next token in a stream of language. It is extraordinarily good at that. The difficulty is that a great many tasks people hand to these models are not, underneath, sequence-prediction over language — and wherever the true object is something else, the model is not weak but blind.
The blind spots are structural, not incidental, and they recur in predictable places: exact arithmetic, where digits split across token boundaries and carries fail; two-dimensional structure, where a spreadsheet is flattened into a line and the model loses which value sits in which column; deep multi-step logic, where accuracy falls off a cliff as the problem deepens; and the model's own output, which a single forward pass has no way to check before it is written.
A fluent, confident, wrong answer is the characteristic result. Not because the model is careless, but because arbitrary-precision arithmetic is simply not what next-token prediction computes.
One failure that is really two
Rigour matters most where a result is widely repeated. A well-known 2024 study showed that adding a single irrelevant clause to a maths word problem could collapse a model's accuracy — evidence, it was argued, that models pattern-match rather than reason. That result is often still cited as though it described current systems. It largely does not: a 2026 replication on frontier models found that once the distractors are audited down to genuinely irrelevant ones, the drop is indistinguishable from zero. The models were making the reasonable judgement that added information might be relevant.
But a second, adjacent failure has not gone away. As the compositional depth of a problem grows, accuracy collapses past a threshold — and, tellingly, spending more inference-time compute does not rescue it; near the collapse point the models reduce their own effort. The lesson is not that models cannot reason, but that the two failures are different and must be handled differently.
The fix is architectural
For exact work, the model should hand the computation to a deterministic tool and confine itself to setting up the problem and explaining the result. For deep combinatorial logic, it should translate the problem into a formal object and let a solver do the search. For structure, the data should be held in a store and queried, not pasted into the prompt. And for anything it generates, a checker — often a piece of running code, not another model — should confirm what cannot be self-checked.
The full field guide catalogues each blind spot, grades the evidence behind it, and gives the routing rule that resolves it.
Full paper→Read the full field guideReferences
- Mirzadeh et al. (2025). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in LLMs. ICLR 2025, arXiv:2410.05229.
- Sturgeon / LessWrong (2026). Revisiting GSM-Symbolic on 2026 frontier models — audited distractor drop indistinguishable from zero.
- Shojaee et al. (2025). The Illusion of Thinking — performance collapse beyond a compositional-depth threshold.
- Lawsen (2025). The Illusion of the Illusion of Thinking. arXiv:2506.09250.
- Wu, Ritter & Xu (2025). Tabular Data Understanding with LLMs: A Survey. arXiv:2508.00217.