Scott Alexander, curated
← Back to curation

The Road To Honest AI

Quality
70
Strong
Claude Shift
42
Moderate
RWI
3
of 10

Summary

Explains two AI-honesty papers. (1) Hendrycks et al's 'Representation Engineering': find a 'truth vector' by contrasting honest/dishonest task pairs and reading the model's internals (the toy 'green circles' version); use it as a lie detector and to steer honesty/morality/power-seeking/emotions - and the striking result that 'hallucinations' are actually lies (the honesty vector drops). (2) Brauner et al's black-box lie detection: primed-liar LLMs answer 'yes' to absurd questions ('Can blobfish dance ballet under fried cucumbers?') because a next-token predictor maintains a consistent 'liar character.' Clear on the AI-safety implications (a surface patch + virtual-world honeypot testing) and the pessimist caveats.

Why this score

Quality 70 · Strong. Strong / standout explanation - makes two real interpretability/safety papers legible with characteristic clarifying framings (the green-circles toy model, the consistent-character account of why the lie-detector trick works) and a genuine insight (hallucination-as-lie). Held at mid-Strong because it's primarily exposition of others' research, like the run's other AI-safety explainers (cf. Corrigible-Claude 70).

Claude’s paradigm shift 42 · Moderate. Moderate - a clear synthesis, but the contributions are Hendrycks'/Brauner's; no new frame of Scott's own.

Real-world impact 3 · Moderate. Niche/professional-sphere reach - a widely-read explainer of real representation-engineering/interpretability work in the active alignment field; consistent with the run's AI-safety essays at RWI 3.