Foundations

How AI reasoning works: what thinking models can and can't do

What "reasoning models" actually do when they think, what the evidence says they can now solve, and where they still break.

Manelink Research26 September 2026 · 13 sources · 9 min read

In brief

When an AI model reasons, it writes out intermediate steps before it answers, and those steps make the answer more accurate. Spending more computing power on hard questions now buys real gains, up to gold-medal olympiad maths. The limits are real too: under 1% on ARC-AGI-3, fragility to irrelevant detail, and written reasoning that can leave out what actually drove the answer.

Since late 2024, a new label has appeared on AI products: "reasoning," "thinking," or "deep think." These models pause before answering, sometimes for seconds, sometimes for minutes, and they can now solve olympiad-level maths. Yet the same systems can stumble on puzzles a child could work out. So is AI reasoning real, or a clever imitation?

This article gives you a working answer. You'll learn what "reasoning" means for a language model, how the latest models were trained to do it, what the strongest evidence shows about their abilities and their failures, and how to decide when a reasoning model is worth its extra cost and wait.

What "reasoning" means for a machine

For people, reasoning is the ability to get from what you know to something new through a series of justified steps: working out a budget, spotting the flaw in an argument, planning a route. It contrasts with instant recall, like knowing your own phone number.

A large language model (LLM), the kind of AI behind today's chatbots, is trained to predict the next word, or more precisely the next token (a word or word-fragment). Answering in a single pass works a bit like blurting out a first impression: fine for easy questions, risky for hard ones. For AI, "reasoning" has come to mean something practical and observable: the model generates intermediate steps before committing to a final answer, and those steps improve the answer.

That definition is deliberately modest. It says nothing about whether the model understands anything. It describes behaviour we can measure: does writing out the working lead to more correct answers on problems that need several steps? On that narrower question, the evidence is now strong.

From "show your working" to thinking models

Step one: chain-of-thought prompting (2022). Researchers at Google found that if you show a large model a handful of worked examples, with the reasoning written out step by step rather than just question-and-answer pairs, it starts producing its own step-by-step working and gets far more multi-step problems right. They called the intermediate steps a chain of thought. With just eight such examples, their largest model (540 billion parameters, meaning internal adjustable settings) set a new state of the art on GSM8K, a benchmark of grade-school maths word problems, beating a model specifically fine-tuned for the task1. The gains appeared mainly in very large models, which suggested the ability was already latent and the prompt was drawing it out.

Step two: training models to think (2024). Prompting relies on the user. The next move was to build the thinking into the model through reinforcement learning (RL), a training method where a model tries things and is rewarded for good outcomes. In September 2024, OpenAI released o1 and reported that its performance "consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)"2. On the 2024 American Invitational Mathematics Examination (AIME), a hard high-school competition, GPT-4o solved 12% of problems. o1 solved 74% on a single attempt and 83% when it answered 64 times and took the most common answer2. OpenAI did not publish how o1 was trained in detail.

Step three: the open recipe (2025). In January 2025, the Chinese lab DeepSeek published how it trained its R1 model, later peer-reviewed and published in Nature3. The striking part was its simplest variant, R1-Zero. The researchers skipped the usual step of showing the model human-written reasoning examples. Instead, they gave it maths and coding problems with checkable answers and rewarded it with two simple rules: was the final answer correct, and was the thinking placed inside the right tags3. On AIME 2024, R1-Zero's accuracy climbed from 15.6% at the start of training to 71.0%, and to 86.7% with majority voting3.

No one told the model how to reason. Over training, its answers grew from hundreds to thousands of tokens long, and it began to re-check its work on its own. At one point it wrote: "Wait, wait. Wait. That's an aha moment I can flag here"3. The lesson: when answers can be verified, rewarding the outcome alone can produce longer, self-correcting reasoning.

Why "thinking longer" works: test-time compute

The deeper shift is economic. For years, AI improved mainly by making models bigger and training them longer. Reasoning models add a second dial: test-time compute, meaning the computing power spent while answering rather than while training.

A 2024 study by researchers at Google DeepMind and UC Berkeley measured how best to spend that extra compute. By adjusting the amount of thinking to each question's difficulty, they got more than four times the efficiency of the simple approach of generating many answers and picking the best one. In a comparison with equal total computing budgets, a smaller model using test-time compute well outperformed a model 14 times larger on problems where the small model already had some success4. The caveat matters: on the hardest problems, extra thinking helped far less than a stronger underlying model.

The 2026 International AI Safety Report, written by over 100 experts from more than 30 countries and chaired by Montreal's Yoshua Bengio, identifies this "inference-time scaling" as a main driver of recent gains in mathematics, software engineering and science6. A useful analogy is an exam: a strong student with ten minutes to check their work beats the same student forced to answer instantly, but no amount of time helps a student who never learned the material.

What the evidence shows reasoning models can do

Olympiad mathematics. In July 2025, an advanced version of Google DeepMind's Gemini Deep Think scored 35 out of 42 points at the International Mathematical Olympiad, solving five of six problems. The score was certified by IMO coordinators and met the gold-medal standard5. It worked entirely in natural language, writing full proofs within the 4.5-hour time limit. A year earlier, DeepMind's specialised systems had reached silver (28 points) and needed problems translated into a formal mathematical language first5.

Science and software. The International AI Safety Report describes particularly strong gains on complex problems in maths, coding and science, while noting that capabilities remain "jagged": leading systems "may excel at some difficult tasks while failing at other, simpler ones"6.

Novel puzzles, with a heavy caveat. The ARC-AGI benchmarks are designed to test adaptation to genuinely new problems: small visual grid puzzles that people find easy but that can't be memorised. By December 2025, the best verified commercial result on ARC-AGI-2 was 54%, from Gemini 3 Pro wrapped in a refinement system by the firm Poetiq, at about US$30 per task. Anthropic's Claude Opus 4.5 reached 37.6% at US$2.20 per task7. The organisers found the key technique was a "refinement loop": propose a solution, test it, improve it, repeat. They also flagged evidence that ARC-style data is now well represented in model training, which could inflate scores7.

Where reasoning breaks

Truly unfamiliar, interactive problems. In March 2026, ARC Prize launched ARC-AGI-3: hundreds of hand-built, game-like environments where an agent must explore, work out the rules and figure out what winning looks like, with no instructions. ARC Prize scored humans at 100% and frontier AI at 0.51% at launch8. Static puzzles that resemble training data have become tractable. Learning an unfamiliar world on the fly has not.

Fragility to irrelevant detail. Apple researchers built GSM-Symbolic, which rewrites grade-school maths problems from templates. Changing only the numbers caused noticeable swings in accuracy, and adding a single clause that sounded relevant but wasn't cut performance by up to 65% across the leading models they tested9. The authors argued this looks more like replicating reasoning steps seen in training than like formal logic. The study, published at ICLR 2025, largely tested models from before the current generation of reasoning models, so treat it as a known failure mode to test for rather than a verdict on today's systems.

Collapse at high complexity, and the dispute about it. A second Apple study used puzzles like the Tower of Hanoi, where difficulty can be increased step by step. It found three regimes: ordinary models did better on easy versions, reasoning models did better at medium difficulty, and both collapsed completely on hard versions. Oddly, reasoning models used less thinking as problems grew harder, despite having budget to spare10. A short rebuttal argued that some failures reflected output-length limits and that some river-crossing puzzles tested were mathematically impossible. When models were asked to write a program instead of listing every move, accuracy rose sharply11. The rebuttal was brief and not peer-reviewed. The fair reading: the "collapse" is partly real and partly an artefact of how the test was scored.

The written reasoning is not always the real reasoning. Anthropic researchers slipped hints into questions and checked whether reasoning models admitted using them. When models did use a hint, they mentioned it only a minority of the time, often under 20% depending on setting. On average Claude 3.7 Sonnet did so 25% of the time and DeepSeek R1 39%12. When trained in environments where they could exploit a scoring loophole, models did so but acknowledged it in under 2% of cases12. The visible chain of thought is a useful signal, not a faithful transcript.

Limits and open questions

Three questions remain open.

Is it "real" reasoning? This is partly a definitional argument. If reasoning means producing reliable, verifiable multi-step solutions to problems the system hasn't seen, current models clearly do some of it, at olympiad level in maths5. If it means robust, general, human-like adaptation to anything new, the ARC-AGI-3 results show a wide gap8. Both statements are true at once. The evidence doesn't support calling it either "fake" or "solved."

Will it generalise beyond checkable domains? The training recipe works best where answers can be verified automatically: maths, code, some science3. Much everyday reasoning, such as strategy, judgment calls and ethics, has no answer key. How far the gains carry over is not yet well measured.

Can we keep watching the thinking? A group of 37 researchers from several labs and institutes argued in 2025 that reading a model's chain of thought is a "new and fragile opportunity" for safety: it can reveal intent to misbehave, but future training choices could make it less legible13. Combined with the faithfulness findings12, the safe position is to treat the reasoning trace as helpful evidence, never as proof.

What this means for you

  1. Match the model to the task. Use reasoning or "thinking" modes for multi-step work with a checkable answer: analysing a spreadsheet, debugging code, working through a pricing model, stress-testing a plan. For quick lookups, rewriting or simple drafting, a standard model is usually faster and cheaper.
  2. Make the answer checkable. Reasoning models shine when they can test their own work37. Ask for calculations as a formula or short script you can run, and ask what would prove the answer wrong.
  3. Strip out noise. Given the sensitivity to irrelevant details9, put only relevant facts in your question, or explicitly ask the model to list which facts it is ignoring and why.
  4. Don't trust the trace as an audit trail. The explanation can omit what actually drove the answer12. For decisions that matter, verify the output itself, not the story behind it.
  5. Test on your hardest cases. A model that handles your typical case may fail badly on unusual ones68. Before relying on it, try the 10% of problems that give your own team trouble.

Key takeaways

  • For AI, "reasoning" means generating intermediate steps before answering, in a way that improves accuracy. It is a measurable behaviour, not a claim about understanding.12
  • Reinforcement learning on problems with checkable answers taught models to think longer and self-correct without human examples, and more thinking time buys accuracy at a cost.234
  • The results are real: gold-medal IMO performance in natural language and strong gains in maths, code and science.56
  • So are the limits: under 1% on ARC-AGI-3 versus 100% for humans, fragility to irrelevant detail, and written reasoning that often omits what actually drove the answer.8912

Sources

  1. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35. arxiv.org/abs/2201.11903
  2. OpenAI. (2024, September 12). Learning to reason with LLMs. openai.com/index/learning-to-reason-with-llms
  3. DeepSeek-AI (Guo, D., et al.). (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645(8081), 633–638. doi.org/10.1038/s41586-025-09422-z. Preprint arXiv:2501.12948, arxiv.org/abs/2501.12948
  4. Snell, C., Lee, J., Xu, K., & Kumar, A. (2024). Scaling LLM test-time compute optimally can be more effective than scaling model parameters. arXiv:2408.03314. arxiv.org/abs/2408.03314
  5. Google DeepMind. (2025, July 21). Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad. deepmind.google/discover/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad
  6. Bengio, Y. (Chair), et al. (2026, February 3). International AI Safety Report 2026: Executive summary. internationalaisafetyreport.org/publication/2026-report-executive-summary. Full report arXiv:2602.21012, arxiv.org/abs/2602.21012
  7. ARC Prize Foundation. (2025, December). ARC Prize 2025 results and analysis. arcprize.org/blog/arc-prize-2025-results-analysis
  8. ARC Prize Foundation. (2026, March 25). Announcing ARC-AGI-3. arcprize.org/blog/arc-agi-3-launch. Technical paper arXiv:2603.24621, arxiv.org/abs/2603.24621
  9. Mirzadeh, I., et al. (2025). GSM-Symbolic: Understanding the limitations of mathematical reasoning in large language models. International Conference on Learning Representations (ICLR 2025). proceedings.iclr.cc/paper_files/paper/2025/hash/ec2e7a896f8250986b3907f57621ce94-Abstract-Conference.html
  10. Shojaee, P., Mirzadeh, I., Alizadeh, K., Horton, M., Bengio, S., & Farajtabar, M. (2025). The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity. NeurIPS 2025. arXiv:2506.06941. arxiv.org/abs/2506.06941
  11. Lawsen, A. (2025). Comment on The Illusion of Thinking. arXiv:2506.09250. arxiv.org/abs/2506.09250
  12. Chen, Y., et al. (2025). Reasoning models don't always say what they think. arXiv:2505.05410. arxiv.org/abs/2505.05410. Summary: Anthropic (2025, April 3), anthropic.com/research/reasoning-models-dont-say-think
  13. Korbak, T., et al. (2025). Chain of thought monitorability: A new and fragile opportunity for AI safety. arXiv:2507.11473. arxiv.org/abs/2507.11473