Research Note
Research Note: E072 Paper Version and Publication Boundary
E072 was recorded after version 1 of *The Illusion of Thinking* appeared on June 7, 2025. That is the paper Dalton read and criticized. The current arXiv record is versio
In this article
Research Note: E072 Paper Version and Publication Boundary
The version problem
E072 was recorded after version 1 of The Illusion of Thinking appeared on June 7, 2025. That is the paper Dalton read and criticized. The current arXiv record is version 3, posted November 20, 2025, and identified as the NeurIPS 2025 camera-ready version with additional discussion in the appendix.
The public package must not silently use the current paper as though Dalton had read it in June. The Episode Story should describe his reaction to version 1, then explain what later versions and responses added. The evergreen paper explainer should use version 3 as the current research record and visibly identify the version.
What the current paper preserves
The current paper still reports three complexity regimes across controlled puzzle environments. Standard non-thinking models can perform as well as or better than reasoning variants on lower-complexity tasks. Reasoning variants gain an advantage in an intermediate range. Both categories then approach zero accuracy on higher-complexity configurations in the reported experiments.
The paper also preserves its claim that reasoning effort initially increases and later falls as complexity rises, even when the allowed generation budget has not been exhausted. Its appendix now directly answers objections about output length, sampling, River Crossing solvability, explicit algorithms, and tool use.
The authors narrowed the Tower of Hanoi range by removing cases above twelve disks. They also refocused River Crossing analysis on configurations below six pairs after acknowledging that puzzle dynamics change at larger sizes. They maintain that the earlier failure at three pairs remains evidence of a planning and constraint-satisfaction limitation.
What later comments establish
The comments do not form a single rebuttal.
Lawsen's June 2025 comment argues that output format and impossible River Crossing instances materially distorted the apparent collapse. Its alternative Tower of Hanoi representation used generated code rather than a fully enumerated move list. The paper describes this as preliminary testing rather than a high-powered replication.
Khan, Madhavan, and Natarajan argue that static text-only evaluation confounds model reasoning with execution under a restricted interface. Their "agentic gap" is a system-boundary argument. It does not prove that tools erase every reasoning failure.
Dellibarda Varela and coauthors tested stepwise and agentic variants. They found that Tower of Hanoi failures were not purely an output-limit artifact and still appeared around eight disks in their setup. They also found that River Crossing performance changes sharply when tests are restricted to solvable configurations. Their conclusion rejects both a universal reasoning-collapse story and a universal rescue-by-tools story.
Publication boundary
The safe public conclusion is narrow. The original paper demonstrated important failures for named models, tasks, prompts, budgets, and scoring rules. Some broad interpretations depended on task validity, representation, and the chosen unit of analysis. Later work qualified parts of the record but did not settle what machine reasoning is or establish that a production system will succeed.
The episode's statement that Apple published the paper as a corporate "cop-out" is a motive inference with no supporting evidence in the paper. It may be preserved as Dalton's June 2025 opinion, followed by an explicit correction. It must not be presented as fact.
Existing public route
The Ghost URL redirects to:
https://www.daltonanderson.net/venture-step/apples-ai-strategy-the-flawed-illusion-of-thinking/
The canonical route, Ghost route, Spotify episode, and YouTube recording returned HTTP 200 on July 28, 2026. The replacement Episode Story should preserve the slug, identify the original publication date, add a visible revision note, and link to the evergreen paper and system explainers. Publication remains unauthorized.
Sources
Follow the evidence.
- machinelearning.apple.com: illusion of thinkingmachinelearning.apple.com
- NIST AI RMF Measure guidanceairc.nist.gov
- youtu.be: 2unoT550UWAyoutu.be
- arxiv.org: 2507arxiv.org
- openai.com: a practical guide to building ai agentsopenai.com
- arxiv.org: 2506arxiv.org
- anthropic.com: demystifying evals for ai agentsanthropic.com
- arxiv.org: 2506arxiv.org
- open.spotify.com: 23EBn3y6n59SKM2C6T2lDsopen.spotify.com
- arxiv.org: 2506arxiv.org
- arxiv.org: 2506arxiv.org
- daltonanderson.ghost.io: apples ai strategy the flawed illusion of thinkingdaltonanderson.ghost.io
- anthropic.com: building effective agentsanthropic.com