Scott Alexander, curated
← Back to curation

Gwern’s AI-Generated Poetry

Quality
68
Strong
Claude Shift
48
Moderate
RWI
2
of 10

Summary

Showcases Gwern's GPT-2 retrained on the Gutenberg Poetry Corpus, correcting Scott's earlier assumption that an AI that never heard sound couldn't do rhyme/meter (it just needed more data). Highlights: flawless iambic pentameter (the pseudo-Homeric Menoetiades passage); rhyme that holds for a couplet then degenerates (the Pope tests — perfect first two lines, then nonsense words like 'wholubil'); occasional brilliance ('Fairy, and wild, and fair, and whole'); the looping 'Emperor Wu (the great Wu), majestical/rapacious' as accidental experimental-political-poetry; and a placebo — an excerpt of Tennyson's 'Ulysses' slipped in to test whether real poems read as fake when you assume they're AI (Scott admits he'd have flagged Tennyson's 'gloom' as a verb-error). The first two lines are always best; he asks why.

Why this score

Quality 68 · Strong. 68 — Strong. An engaging, well-curated AI-poetry showcase (a sequel to the GPT-2 essay) with a genuinely sharp epistemics move — the Tennyson placebo exposing our AI-skepticism bias — plus the first-two-lines-best observation. More showcase than argument, which places it mid-Strong.

Claude’s paradigm shift 48 · Moderate. 48 — Moderate. The placebo-test insight (our priors make us misread poetry as AI-flawed) is a fresh, memorable point; the showcase itself is curation.

Real-world impact 2 · Minor. 2 — a memorable early AI-poetry demonstration; within-discourse reach.