Wire flash
Tsinghua, Oxford, Stanford study: LLMs continue reasoning even when told to disable thinking
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
A new paper from Tsinghua, Oxford, and Stanford, titled 'Thinking Inertia: LLMs Keep Thinking When Told Not To,' finds that large language models continue to output reasoning even when explicitly instructed to disable thinking. The study reports that with thinking disabled, DeepSeek-V4-Flash still wrote out reasoning in 99.9% of open-ended answers. The researchers note that a missing think tag or a short reply does not prove the model skipped reasoning. Yes/no questions are easiest for models to answer directly, while open-ended questions strongly pull models back into reasoning. Forcing answer-only replies on open-ended tasks raised compliance to about 40% across 5 models but reduced accuracy by approximately 15 points, according to the paper.
Source report
A new paper from researchers at Tsinghua University, Oxford, and Stanford reveals that large language models (LLMs) continue to reason out loud even when explicitly instructed to stop thinking—particularly on open-ended questions.
Key Findings
- Persistent reasoning: With thinking disabled, DeepSeek-V4-Flash still produced reasoning in 99.9% of open-ended answers.
- Misleading indicators: A missing think tag or a short reply does not prove the model skipped reasoning.
- Question type matters: Yes/no questions are easiest for models to answer directly; multiple-choice questions fall in the middle; open-ended questions consistently pull models back into reasoning.
- Trade-off with forced compliance: Forcing answer-only replies on open-ended tasks raised compliance to about 40% across five models, but cut accuracy by approximately 15 points.
Source
- Paper: "Thinking Inertia: LLMs Keep Thinking When Told Not To"
- Available at: arxiv.org/abs/2610.11765
Source
rohanpaul_aiNeutral / independent