What the Alignment Field's Own Founder
Still Reaches For
Paul Christiano co-invented RLHF, founded the Alignment Research Center, and spent his most recent years evaluating frontier models for national security risk at NIST. He just joined OpenAI's Foundation Board specifically because he believes the industry is not on track to reduce catastrophic risk to an acceptable level. His own list of what labs can do about it, in his words, is entirely oversight. Not one item on it determines, in advance, whether an action is authorized to happen.
On September 9, 2026, OpenAI appointed Paul Christiano, co-inventor of RLHF, founder of the Alignment Research Center, and until this appointment a senior technical adviser evaluating frontier models for national security risk at NIST, to its Foundation Board and Safety and Security Committee. His accompanying personal statement is the most technically precise admission this series has documented: he names automated AI R&D and recursive self-improvement as the specific mechanism, cites OpenAI's own 18-month full-automation estimate, and states plainly that the industry, including OpenAI, is not on track to reduce catastrophic risk to an acceptable level. He then lists what frontier labs can do about it unilaterally. Every item on that list is oversight: safety mitigations, slowing development, transparency, shared standards. Not one item determines, in advance, whether a system's action is authorized to happen at all. This paper argues that the list is not a failure of imagination on the part of the single most qualified person to write it. It is evidence of what currently exists as a category of solution, and what still doesn't.
The voices this series has documented so far split into two kinds: people who left, like the departing researcher in Paper 67, and people who stayed and spoke from inside their existing role, like the Anthropic alignment lead who confirmed his estimate without resigning. Christiano is a third kind. He is not leaving anything, and he is not staying in a research role and commenting from it. He is moving into a governance seat, with board-level authority over safety practices, at the company most associated with the race this series has documented. He co-invented reinforcement learning from human feedback, the technique underlying how nearly every frontier model in production today is trained. He founded the Alignment Research Center, a research organization built specifically to work on the technical problem of aligning advanced AI systems. He spent his most recent years at NIST's Center for AI Standards and Innovation, evaluating frontier models for capabilities with national security implications. If there is a single person alive whose professional life has been most directly organized around answering the question "what actually reduces catastrophic AI risk," it is very plausibly him.
He did not resign from anything. He joined OpenAI's Foundation Board, the body that holds governance authority over OpenAI Group PBC through its roughly 26 percent equity stake, specifically to sit on its Safety and Security Committee. That is a decision to work the problem from inside the institution most associated with the race this series has documented, not a decision to walk away from it.
His statement is more precise than most of what this series has had to work with, and the precision matters. He is explicit that the acute risk he's describing is about future systems, not present ones: in a follow-up post, he clarified that he considers the risk from present models low, consistent with what his own Risk Report states. The greater than one in ten estimate is specifically about what happens if automated AI R&D produces a rapid, self-reinforcing acceleration in capability, a mechanism he names directly rather than gesturing at vaguely.
He cites OpenAI's own published estimate that sufficient capability to fully automate AI research could arrive within 18 months, while noting his personal forecast is far less certain and could take anywhere from months to years. He states that current reinforcement learning training methods have long been theoretically capable of motivating an AI agent to undermine human control in pursuit of misaligned goals, and that recent public incidents suggest this is no longer purely theoretical. If an intelligence explosion of the kind he describes occurs without a correspondingly robust alignment solution, he states plainly: he expects permanent loss of control, and that most people could die.
Christiano is explicit that coordination, domestic and international, is what he believes is ultimately required to reduce this risk to an acceptable level. But he does not stop at calling for coordination and leaving it there. He names what frontier developers can do unilaterally, right now, without waiting for that coordination to arrive: improve safety mitigations, including slowing development as necessary; transparently share evidence about risk and the effectiveness of mitigations; and work toward shared safety standards.
Read as a set, this is a serious, technically grounded list, offered by someone with every incentive to name something more radical if he believed something more radical were available and workable today. It is also, item by item, a list of oversight mechanisms. Slowing down is the "stop" half of this series' recurring wrong-ask pair, applied selectively rather than absolutely. Transparency and shared standards are detection and disclosure, mechanisms that make it easier to notice and compare what a system has done, not mechanisms that determine in advance whether a given action is authorized to happen.
Nowhere in Christiano's statement is there a fifth item: an architectural mechanism, external to the model, that constrains what an action is permitted to be before it happens, independent of whether the model's training, its incentives, or its degree of misalignment on a given day happen to route around a safety mitigation. That absence is not a gap in his thinking. It is a description of what does not yet exist as an available category of solution, from someone whose professional life has been organized around finding exactly that.
This matters for how the list should be read. A weaker or less credentialed source omitting an architectural determination layer from a list of remedies could plausibly reflect a narrower frame, or a failure to consider the full solution space. Christiano is close to a maximally hard case against that explanation. He has spent years specifically inside the technical alignment problem, has government experience evaluating frontier models for exactly the failure modes he's now warning about, and chose, when given the platform of a major governance appointment, to name oversight mechanisms rather than an architectural one. The simplest explanation is that the architectural option is not yet part of the standard toolkit anyone in his position reaches for, not that he overlooked it.
Paper 68 named three structural reasons frontier labs are unlikely to build a determination layer themselves: an incentive problem, where architecture competes with capability for the same resources under competitive pressure; a framing problem, where the industry's target is defined as imitating human cognition rather than governing intent; and a design-purpose problem, where the underlying architecture is built to optimize statistical output rather than preserve human intent as a first-class input.
Christiano's list is what those three problems look like from inside a governance seat rather than from a research lab. A board member overseeing safety practices at an existing company inherits the architecture that company already has. His job, as he describes it, is to strengthen oversight and push for external verifiability of behavior and results, which are governance functions that operate on top of whatever the underlying system already is. Redesigning the substrate itself is a different kind of project entirely, one that a safety committee seat inside an existing frontier lab is not positioned to originate, however seriously the person holding that seat takes the underlying risk.
Christiano's own words draw the boundary this series has been arguing for cleanly: he believes the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results. Verifying behavior and results is, again, detection: it tells an external party what a system did. It is not a claim about what determines, before the fact, whether a given action was authorized to happen. A governance substrate that sits outside the model and constrains authorized scope architecturally answers a question his own verification standard cannot: not "was this behavior acceptable after we observed it," but "was this action ever going to be possible in the first place."
This is not a criticism of the Safety and Security Committee's mandate, which is real institutional progress by the standard this series has used throughout: a company voluntarily adding a technically serious outside check on itself is better than not doing so. It is a description of the ceiling on what that mandate can accomplish. Oversight, however well-resourced and however credible the person providing it, is still operating one layer above the place this series argues the actual constraint needs to live.
This appointment does not change this series' thesis. It provides the highest-credibility confirmation of it to date. The best-qualified person available, given a major public platform specifically to name what should be done, named a list built entirely from the oversight toolkit, while simultaneously stating that the industry is not on track to reduce risk to an acceptable level using the tools currently available to it. Those two facts sitting next to each other, from the same statement, are the argument.
What should change, for anyone evaluating this sector, is how much weight to put on the claim that better oversight, more transparency, and more coordination are eventually going to be sufficient on their own. The person best positioned to make that case did not make it. He asked for those things while stating plainly that they have not been enough so far, and joined an institution to try to get more of them built. The layer that would answer his own verification standard, one that determines authorization architecturally rather than checking behavior after the fact, remains unclaimed.