Highlights From The Comments On AI Geoguessr
Read the original on Astral Codex Ten →
Follows up on
↳ Testing AI's GeoGuessr Genius — Essay · May 2025
Summary
Companion to the AI-GeoGuessr post. Collates readers' own o3 tests (mixed; US/Anglo bias), the 'good at touristy spots' rule of thumb, the Everest-guess-was-actually-impressive analysis, the Trevor-Rainbolt human-GeoGuessr comparison, and Daniel Kang's key point: o3 was almost certainly trained on GeoGuessr-style verifiable tasks, so this shouldn't update you much on general intelligence.
Why this score
Quality 67 · Strong. Companion band: a genuine analytical payoff (the don't-over-update framing + the touristy-spots rule), held below the 71 parent (ACX-697).
Claude’s paradigm shift 46 · Moderate. Notable: the trained-on-verifiable-tasks caveat reframes the capability demo.
Real-world impact 1 · Negligible. Within-discourse companion; no material effect. =1.