Don’t be fooled—LLMs don’t reason
- Published
- Oct 2, 2026 — 08:00 UTC
The article discusses the limitations of large language models (LLMs) in reasoning capabilities, drawing comparisons to historical AI achievements such as AlphaGo and Deep Blue. It emphasizes that while LLMs can generate human-like text, they do not possess true reasoning abilities.
The piece references AlphaGo, developed by DeepMind, which famously defeated professional Go player Lee Sedol in March 2016 with a score of 4-1. This match is often cited as a landmark moment in AI, showcasing the program's ability to make strategic decisions that appeared creative. Lee Sedol himself remarked on the unexpected nature of AlphaGo's play, particularly highlighting a pivotal move (Move 37) that had a mere 1 in 10,000 probability of being executed by an expert human player. This led him to reconsider the nature of the AI, suggesting that it exhibited a form of creativity.
Thore Graepel, a core member of the AlphaGo team and chair of machine learning at University College London, commented on the significance of Move 37, stating that it demonstrated how a machine could evaluate potential future positions and make decisions that might defy conventional human instincts. This contrasts sharply with the capabilities of LLMs, which, despite their impressive text generation, lack the underlying reasoning processes that characterize strategic decision-making in games like Go or chess.
The article also references Deep Blue, the chess-playing computer developed by IBM that defeated world champion Garry Kasparov in 1997. Deep Blue's ability to evaluate 200 million chess positions per second exemplifies a different approach to AI, one that relies on brute computational power and strategic evaluation rather than the probabilistic text generation seen in LLMs.
Overall, the article argues that while LLMs can mimic reasoning through language, they do not engage in true reasoning processes akin to those demonstrated by earlier AI systems like AlphaGo and Deep Blue. This distinction is crucial for understanding the current limitations of LLMs in tasks requiring genuine cognitive reasoning.
By Turing Wire Research Desk · Oct 2, 2026 · How we work →
Summarised from MIT Technology Review's coverage by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: MIT Technology Review
