Scott Alexander, curated
← Back to curation

Highlights From The Comments On AI Geoguessr

Quality
67
Strong
Claude Shift
46
Moderate
RWI
1
of 10

Follows up on

Testing AI's GeoGuessr Genius — Essay · May 2025

Summary

Companion to the AI-GeoGuessr post. Collates readers' own o3 tests (mixed; US/Anglo bias), the 'good at touristy spots' rule of thumb, the Everest-guess-was-actually-impressive analysis, the Trevor-Rainbolt human-GeoGuessr comparison, and Daniel Kang's key point: o3 was almost certainly trained on GeoGuessr-style verifiable tasks, so this shouldn't update you much on general intelligence.

Why this score

Quality 67 · Strong. Companion band: a genuine analytical payoff (the don't-over-update framing + the touristy-spots rule), held below the 71 parent (ACX-697).

Claude’s paradigm shift 46 · Moderate. Notable: the trained-on-verifiable-tasks caveat reframes the capability demo.

Real-world impact 1 · Negligible. Within-discourse companion; no material effect. =1.