Scott Alexander, curated
← Back to curation

Book Review: Human Compatible

Quality
78
Excellent
Claude Shift
60
Notable shift
RWI
4
of 10

Summary

A substantive review of Stuart Russell's Human Compatible (tagged book review/lw). Three throughlines. (1) The book as artifact: a top AI scientist (Berkeley CS, leading textbook author) writing a calm, respectable, dad-joke-laden case for AI safety that 'somehow makes risk from superintelligence not sound weird' — a better credentialing artifact than Bostrom's deliberately strange Superintelligence. (2) Russell's 'Not-So-Great Debate' chapter dismantling skeptics (Kelly's 'intelligence isn't one-dimensional,' Ng's 'overpopulation on Mars,' AI100's 'superintelligence is impossible') via proving-too-much and common sense — Scott's clearest case that the bar for the basic worry is low. (3) Russell's research program: Cooperative Inverse Reinforcement Learning recast as 'treat the command as evidence of a mysterious true goal,' which dissolves corrigibility (the AI accepts being switched off as information) and wireheading — one of the better lay explanations of value-learning. The standout original contribution is the Misuses-of-AI critique: Scott pushes back (with citations) on deepfake / algorithmic-bias / radicalization / automation panics as overblown, then raises the prescient worry that AI-safety advocates who endorse weak near-term panics risk 'crying wolf' and losing credibility before crunch time — before conceding society doesn't actually punish being wrong.

Why this score

Quality 78 · Excellent. Excellent, low (78). A top-quartile review that doubles as a clear CIRL primer and adds two genuinely original framings — the 'book as artifact' observation and the 'cry-wolf credibility' worry — both since widely cited. On a par with the canonical AI-safety primers (cf. No-Time-Like-The-Present 77), edging just above on originality of meta-commentary.

Claude’s paradigm shift 60 · Notable shift. Notable, low (60). The technical core (CIRL) is Russell's, not Scott's, so it is not high novelty for the reviewer; but the artifact framing and the cry-wolf credibility argument were fresh, non-obvious contributions to AI-safety discourse in early 2020.

Real-world impact 4 · Moderate. Moderate (4). A notable input to AI-safety / EA-communications discourse — the cry-wolf concern in particular recurs in that subculture — but the consequential real-world change runs through the field, not this post, and the field was still niche in 2020.