// MIT TECH REVIEW — INTELLIGENZA ARTIFICIALE
Don’t be fooled—LLMs don’t reason
Ten years after AlphaGo’s match against Go champion Lee Sedol, today’s AI still isn’t tapping into the machinery that made that win possible.
On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked so absurd that some commentators thought it was a programming glitch. It wasn’t. AlphaGo won the game, ultimately triumphing 4-1 over Lee Sedol, one of the greatest professional Go players of all time. “I thought AlphaGo was based on probability calculation and that it was merely a machine,” Lee said afterwards. “But when I saw this move, I changed my mind. Surely, AlphaGo is creative.”
When Deep Blue defeated then reigning world chess champion Garry Kasparov in 1997, it did so by looking six to eight moves ahead per player and evaluating 200 million chess positions per second, using rules hard-coded by humans. Go is a vastly more complex game. A stone’s worth depends on how distant groups and territory unfold over dozens of moves. Computing even a fraction of the possible outcomes would take a supercomputer billions of years. To win, AlphaGo had to sense who was ahead at a glance and even invent moves no human had thought to play.
That is why many accounts of AlphaGo’s match against Lee portray move 37 as a flash of pure machine intuition. But that is a misunderstanding. It was actually AlphaGo’s powers of reasoning that made this creative choice—and these are powers that today’s AI lacks. If we want future AI systems to produce trustworthy results and really novel insights in fields like science and medicine, we need to equip them with genuine reasoning capabilities of this kind.
AlphaGo is made up of two systems. The first, its policy network, was trained to guess what move a strong human would play. This “intuitive” part regarded move 37 as nothing special—a play that had a roughly one in 10,000 chance of being made by an expert human player. What made AlphaGo choose it was the program’s search machinery, which looked beyond immediate plausibility and weighed the future consequences of proposed moves. It explicitly constructed and searched a game tree with thousands of branches, each representing a different possible future.
A well-known theory in the behavioral sciences, popularized by Daniel Kahneman, distinguishes between two modes of human thought: System 1 is fast, gut-level, effortless; system 2, slow, step-by-step, and deliberative. AlphaGo offered a striking machine analogue of that split. Its networks supplied the hunches—this move looks promising, this position looks won—and its search supplied the deliberation, testing those hunches against the moves and countermoves that would follow. As in human cognition, neither half works alone. Intuition alone would never have opted for move 37, and brute-force search would have struggled to sieve through all the many possible moves.
This is strikingly different from the way today’s AI models work. A large language model picks the next token, over and over. That amounts to system 1 in action—fast, associative, and surprisingly good pattern completion across almost every subject people write about.
Not long after ChatGPT debuted, the field realized that language fluency alone falls short of true usefulness. The apparent solution was to make models that deliberate: Instead of answering immediately, they can now generate intermediate steps that decompose a problem, carry forward partial results, and influence subsequent reasoning—a process known as chain of thought. The gains have proved real, above all in mathematics and coding. But unlike AlphaGo’s search, this does not introduce a genuinely separate reasoning mechanism: The intermediate reasoning is still produced by the same next-token prediction process, iterated for longer before the model commits to an answer.
Three shortcomings prevent what chatbots do fro