TheVortiq
Inteligencia Artificial

Is this the end of LLMs? François Chollet questions the path to AGI

The co-founder of the ARC Prize argues that OpenAI's obsession with language models is setting back the development of true artificial general intelligence by a decade.

October 5, 2026 · 4 min read

Abstract geometric pattern of illuminated lights

TL;DR: François Chollet argues that LLMs are a dead end for AGI. Their focus on data scaling has diverted attention from researching architectures capable of real reasoning, delaying technological progress by a decade.

The stagnation of Artificial General Intelligence: A critical analysis

The global tech industry is immersed in an unprecedented arms race, where the massive scaling of Large Language Models (LLMs) has become the gold standard for measuring progress. However, this hegemony of the token-prediction paradigm is being challenged by one of the sector's most authoritative voices: François Chollet, a software engineer at Google and creator of Keras. Chollet argues that we are facing an architectural dead end: by prioritizing statistical interpolation over true reasoning capability, major labs have diverted critical resources, delaying progress toward Artificial General Intelligence (AGI) by five to ten years.

Why are LLMs not AGI? The myth of generalization

Chollet's critique, laid out extensively on the Dwarkesh Podcast, is based on the ontological distinction between statistical memorization and true generalization. While AGI is defined by the ability to acquire new skills from limited data—an intrinsic characteristic of human biological intelligence—modern LLMs operate under the 'flat learning curve' paradigm. These models depend on massive volumes of pre-existing data to approximate useful results through probability.

Historically, AI has gone through various 'winters' and 'springs.' During the 80s, symbolic AI (GOFAI) promised pure logical reasoning but failed when faced with the uncertainty of the real world. Today, the pendulum has swung to the opposite extreme: massive connectionism. Chollet argues that current models are excellent language simulators, but lack an architecture capable of reasoning in unknown environments. Unlike a human, who can learn to use a new tool after seeing it once, an LLM requires prior training on millions of similar examples to replicate a behavior. This lack of 'data efficiency' is, according to Chollet, empirical proof that we are not looking at general intelligence, but at hyper-optimized probabilistic search engines.

The trap of marketing and biased benchmarks

One of the most controversial points is the manipulation of success metrics. The industry has created benchmarks that largely measure information retrieval (memorization) rather than the ability to solve novel problems. To democratize and objectify evaluation, Chollet promoted the ARC-AGI (Abstraction and Reasoning Corpus). This benchmark is radically different: it proposes tasks that require abstract and logical reasoning, designed so that a system cannot simply 'guess' the answer based on its training corpus.

To date, no LLM has surpassed ARC-AGI with a score close to human level. This stagnation in a benchmark designed to measure learning capacity proves that the multi-billion dollar investment in GPU infrastructure and data centers is producing 'productivity AI' (useful for drafting emails or summarizing documents), but not 'reasoning AI.' The industry, pressured by the need to justify exorbitant market valuations before future IPOs of giants like OpenAI or Anthropic, has turned the narrative of imminent AGI into a marketing tool, hiding structural technical limitations.

Consequences for the tech ecosystem and capital

If Chollet's thesis is confirmed, the implications for the market will be profound and disruptive:

  • Capital revaluation: Current capital allocation is heavily biased toward scaling (more compute, more data). If AGI requires an architectural paradigm shift, much of the current GPU infrastructure could become obsolete or insufficient for the next evolutionary leap.
  • Pivot toward neuro-symbolic AI: AI labs could be forced to abandon the 'Transformer-only architecture' model to integrate symbolic reasoning systems, a field that combines the learning capacity of neural networks with the logical structure of classic expert systems.
  • Transparency and corporate ethics: The pressure to maintain the hype has led companies to publish exaggerated safety incidents or 'short-term' AGI promises to keep investor interest alive. This phenomenon, similar to the dot-com bubble, could lead to a crisis of confidence if models fail to overcome their current reasoning limits.

In conclusion, although LLMs offer undeniable commercial utility in productivity and automation tasks, the consensus that they are the direct path to AGI is increasingly fragile. Chollet's warning is not an attack on AI, but a call to scientific sanity: if we want to reach true general intelligence, we must stop measuring success by the size of models and start measuring it by the ability to reason under uncertainty. The history of computing teaches us that real revolutions do not always come from scaling what already works, but from daring to change the rules of the game when the current model hits a ceiling.

Keep reading