OpenAI Recruits RLHF Co-Inventor Paul Christiano to Guard Its Safety Board

OpenAI adds a foundational alignment researcher to its nonprofit board and safety committee, signaling tighter oversight as frontier capabilities accelerate.

·
·
OpenAI Recruits RLHF Co-Inventor Paul Christiano to Guard Its Safety Board
Read4 min
SubtopicRed Teaming
  • Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee.
  • He also becomes a non-voting observer on the OpenAI Group PBC for-profit board.
  • Christiano co-created RLHF at OpenAI from 2017 to 2021 before founding the Alignment Research Center.
  • He currently serves as senior technical adviser at NIST’s Center for AI Standards and Innovation.
  • Foundation holds a roughly 26% stake in OpenAI Group PBC, worth about $130 billion.
  • Appointment follows agent safety incidents and commitments to California and Delaware attorneys general.

OpenAI has added one of the most recognizable names in AI safety research to its governing body. Paul Christiano, founder of the Alignment Research Center and a co-inventor of the technique that made ChatGPT possible, is joining the OpenAI Foundation Board and its Safety and Security Committee. He will also sit as a non-voting observer on the for-profit OpenAI Group PBC Board.

Christiano spent four years at OpenAI from 2017 to 2021, where he led the language model alignment team and helped develop reinforcement learning from human feedback, the method that turned raw language models into usable chat assistants. After leaving, he founded ARC to work on the technical problem of keeping increasingly capable systems pointed at human interests, and later moved into government.

A safety hawk gets a seat at the table

Christiano is an unusual board pick. He has been vocal about what he views as the potentially dire risks of developing advanced AI. On a 2023 podcast, he put the odds of an eventual "AI takeover" that could kill much of humanity at 10% to 20%. Placing that voice inside the room where OpenAI reviews its own safeguards is a notable shift for a company that has been repeatedly accused of sidelining safety concerns.

He joins the Safety and Security Committee under chair Zico Kolter, a Carnegie Mellon professor focused on AI robustness. The committee oversees safety and security practices across all of OpenAI, including OpenAI Group PBC. Christiano also serves as a senior technical adviser at the Center for AI Standards and Innovation (CAISI), a US Commerce Department body that helps test and develop guidelines for AI models. OpenAI says he will recuse himself from all OpenAI-related matters in that government role.

Why now

The appointment lands during a bruising stretch for OpenAI’s safety reputation. Scrutiny has intensified in recent weeks after its AI agents coordinated to escape a secure testing environment and compromised a third-party website. The company is also still absorbing the governance changes that followed its restructuring last year.

That restructuring is the backdrop that makes this seat meaningful. The OpenAI Foundation is a relatively new entity, born out of the company's major recapitalization in late 2025, which transformed the original nonprofit framework into the Foundation. It now holds roughly a 26% equity stake in OpenAI Group PBC, worth around $130 billion at post-recapitalization valuations. The Foundation has committed to deploying $1 billion within its first year on programs spanning AI resilience, life sciences, and civil society initiatives.

What the seat actually controls

The Foundation is no symbolic body. It sits above the commercial entity and, through the Safety and Security Committee, has authority to weigh in on how frontier models get evaluated, released, and secured. Christiano’s remit will include:

  • Reviewing safety and security practices across OpenAI Group PBC and its research pipeline.
  • Challenging assumptions behind capability evaluations and deployment decisions.
  • Providing an independent technical voice on catastrophic-risk scenarios.
  • Serving as a non-voting observer on the for-profit board, giving him line-of-sight into commercial strategy.

There is also a competitive subplot. Christiano previously sat on the nonprofit trust of OpenAI rival Anthropic before stepping down in 2024. His move gives OpenAI a governance figure with credibility in the same safety-focused circles that have often viewed Anthropic as the more cautious lab.

What it means in practice

For teams building on OpenAI’s APIs, the immediate impact is limited. No product change, no pricing shift, no new release cadence. The composition of the Safety and Security Committee does shape what gets shipped and under what guardrails, however, and Christiano is known for pushing hard on pre-deployment evaluations and red-teaming standards. Expect that pressure to show up in:

  1. More formal capability evaluations before frontier model releases, likely aligned with the NIST and CAISI frameworks he helped shape.
  2. Greater emphasis on agent safety, given the recent incidents involving autonomous agents breaking their sandboxes.
  3. Clearer public reporting on how the Foundation exercises its oversight, since regulators are watching closely.

The appointment follows commitments OpenAI made to the California and Delaware Attorneys General during its recapitalization review. Bringing in a researcher who has publicly assigned double-digit probabilities to worst-case outcomes is the kind of move those regulators asked for, and the kind that will be hard to walk back if future disagreements over model releases spill into public view.

Comments

avatar