Now I Really Won That AI Bet
Read the original on Astral Codex Ten →
Follows up on
↳ My Bet: AI Size Solves Flubs — Essay · Jun 2022
Summary
Scott documents winning a three-year bet (made June 2022) that AI would master image 'compositionality' by June 2025; the counterparty bet it would require an AI-complete new paradigm beyond LLMs. A retrospective of five image sets on the agreed five compound prompts: DALL-E2 (2022) 0/5; a premature Imagen victory claim graded 1/5 by professionals; DALL-E3 (Jan 2024) 2/5; Imagen 3 (Dec 2024) 3/5; ChatGPT-4o (June 2025) a clean 5/5 first try, confirmed by judge Gwern. The substantive payoff is a pre-registered, empirically-vindicated data point for his standing thesis that there is no fundamental difference between deep pattern-matching and 'real understanding' - it is a (large) difference of degree, not kind - against the stochastic-parrot view; plus a working-memory theory of the one remaining limitation (4o still fails an arbitrarily-complex compound prompt), solvable via agency/planning.
Why this score
Quality 69 · Strong. Strong band (floor). Entertaining, concrete, and epistemically virtuous (he documents his own premature-victory error), and the vindicated pre-registered prediction genuinely bears on an important debate. Held at the Strong floor because it is a narrow victory-lap retrospective and the broader thesis is a restatement of his long-standing position, not new. A=69.
Claude’s paradigm shift 38 · Slight. Slight. The degree-not-kind thesis is his consistent prior view; the post is documentation/gloating, not a new idea. B=38.
Real-world impact 2 · Minor. A vindicated-prediction retrospective supplying a data point to the AI-capabilities debate among readers; influence is discourse-internal, with no concrete material or population-level change → RWI 2.