Wire flash
Chollet: LRMs surpass base LLMs by shifting from transductive to inductive reasoning
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
In a post on X, AI researcher François Chollet argues that the critical distinction between base large language models (LLMs) from 2024 and earlier and modern large reasoning models (LRMs) is not symbolic tool use, but a paradigm shift from transductive reasoning (intuiting the answer to a query) to inductive reasoning (intuiting the program or instructions that produce the answer). He states that LRMs are trained to be inductive and perform test-time induction, which unlocks fluid intelligence. Chollet claims base LLMs have approximately zero fluid intelligence, while LRMs possess substantial levels. He cites benchmark performance on ARC 1, noting that LLMs still score only 10-15% today, even after scaling by a factor of 100,000x from 0%. In contrast, LRMs of the same size or smaller saturated the ARC 1 benchmark in 2025.
Source report
The critical difference between base LLMs (2024 and earlier) and modern LRMs is not symbolic tool use. Instead, it lies in a fundamental paradigm shift:
- Base LLMs operate under a transductive paradigm — they intuit the answer directly from the query.
- Modern LRMs operate under an inductive paradigm — they intuit the program or instructions that produce the answer to the query.
Training and Test-Time Induction
LRMs are trained to be inductive, and they perform test-time induction, meaning they predict a natural language program or reasoning chain at test time. This capability unlocks entirely new abilities — most notably, fluid intelligence.
- Base LLMs, to this day, exhibit approximately zero fluid intelligence.
- LRMs demonstrate substantial levels of fluid intelligence.
Performance on ARC 1
The ARC 1 benchmark (introduced in 2019) highlights this gap:
- Base LLMs still achieve only ~10–15% performance today.
- Scaling base LLMs by a factor of roughly 100,000x improved their score from 0% to 10%.
- In contrast, LRMs of the same size or smaller saturated the ARC 1 benchmark in 2025.
Source
fcholletNeutral / independent