Paul Christiano
AI Alignment Researcher
The researcher who pioneered RLHF — the technique that made ChatGPT possible — then left OpenAI to work on alignment full-time, publicly putting roughly coin-flip odds on catastrophe once AI reaches human level.
Why this score
How the formula works →- +3.0antiPredicts AI catastrophe
- +2.2antiWalked away over AI concerns
- +0.9anti2 anti-AI statements (forceful)
- ±0nuancedHolds a documented two-sided view
- ±0nuancedWorks on AI safety
Weights, statement counts, the ×1.5 multiplier because AI is their life’s work, and the formula’s smoothing are already baked in — the points add up to the score (direction-free evidence pulls toward the middle instead): Firmly Anti-AI · 6.1/10
In his words
Overall, maybe we’re talking about a 50/50 chance of catastrophe shortly after we have systems at the human level.
I think maybe there’s something like a 10-20% chance of AI takeover, [with] many [or] most humans dead.
Biography
Paul Christiano (born 1988) is an American AI researcher whose fingerprints are on both sides of the AI story. At OpenAI he ran the language-model alignment team and co-developed reinforcement learning from human feedback (RLHF) — the technique that turned raw language models into usable assistants and made ChatGPT possible. In 2021 he left to found the Alignment Research Center (ARC), a nonprofit focused on the harder problem: ensuring systems smarter than their evaluators remain honest and controllable.
Coin-flip odds from an insider
What makes Christiano’s warnings land is that they are probabilities, not prophecies. On the Bankless podcast in April 2023 he estimated a 10–20% chance of AI takeover with most humans dead, and around 50/50 odds of catastrophe shortly after systems reach human level — figures far above industry consensus, delivered in the flat register of someone reporting a calculation. ARC’s evaluations arm (now METR) pioneered the practice of testing frontier models for dangerous capabilities before release.
In April 2024 the U.S. government made him head of AI safety at the U.S. AI Safety Institute at NIST — a hire that reportedly triggered internal revolt from staff who considered him an “AI doomer,” and criticism from others who noted the irony of an RLHF inventor policing the products of his own technique. Christiano has never advocated abandoning AI; he argues for slowing down at the threshold, investing heavily in alignment, and building the measurement science that would let anyone know when systems become dangerous.
Quote sources
- Bankless podcast, 2023(decrypt.co)
Sources & further reading
Canonical record: https://battlelines.ai/topic/paul-christiano








