Scott Alexander, curated
← Back to curation

Can This AI Save Teenage Spy Alex Rider From A Terrible Fate?

Quality
75
Excellent
Claude Shift
52
Moderate
RWI
3
of 10

Summary

Vivid explainer of Redwood Research's robust-injury-classifier project (fine-tune GPT on 4,300 Alex Rider fanfics -> human-label violence -> train a classifier -> bounty-hunt adversarial examples -> retrain) as a concrete test of prosaic alignment. The adversarial failures (SEO-spam derails, poetic euphemism) motivate the payoff: prompting an agent for what it will do differs from putting it in the situation (the Generalissimo metaphor), so agentic alignment needs situation-prompts + interpretability + ELK, not just this classifier.

Why this score

Quality 75 · Excellent. Strong / low-Excellent edge — one of the better popular explanations of the adversarial-robustness problem; clear, memorable, does real conceptual work (the Generalissimo agentic-vs-prompted point, the three-things-needed synthesis) and aged well. 75.

Claude’s paradigm shift 52 · Moderate. Notable — a fresh, lucid framing/explainer building on Redwood's and ARC's work.

Real-world impact 3 · Moderate. Influential within AI-safety discourse. 3.