In an interview published September 12, Anthropic CEO Dario Amodei discussed slowing AI development and placing independent evaluators inside AI companies; former researcher Jacob Coxon described his resignation and concerns about the race to build more capable systems. The discussion followed OpenAI chief scientist Jakub Pachocki’s September 6 warning that alignment and monitoring remain unresolved. The debate is about whether safeguards can keep pace with increasingly capable AI.

The latest debate centers on oversight

The segment displays Jacob Coxon’s resignation post and Evan Hubinger’s response, alongside an archived clip of Dario Amodei discussing catastrophic risk.

Coxon announced on September 9 that he had resigned from Anthropic after three years of pretraining research at OpenAI and Anthropic. He accused both companies of racing toward self-improving superintelligence. Evan Hubinger, an Anthropic alignment science lead, publicly endorsed Coxon’s concern; that is a personal assessment, not a shared probability for AI risk.

What OpenAI says about alignment and monitoring

In his September 6 essay, “An Alien Mind,” Pachocki said no laboratory had solved alignment and monitoring sufficiently to continue scaling at maximum speed responsibly for much longer. He argued that the pace of development should be constrained by confidence in safety.

OpenAI distinguishes two alignment problems. Goal alignment is whether a system pursues the objective it was given. Value alignment is whether it acts consistently with human values in unfamiliar or adversarial situations. A system can pursue a stated goal while still behaving poorly in circumstances its designers did not anticipate.

ConceptWhat it meansWhy it matters
Goal alignmentPursuing an assigned objectiveA system may meet its stated goal without behaving appropriately in an unfamiliar situation.
Value alignmentActing consistently with human values in unfamiliar or adversarial situationsThe system must generalize beyond a narrow instruction or familiar setting.
Recursive self-improvementAI increasingly contributes to developing successor AI systemsA fully autonomous cycle remains theoretical; no frontier lab has claimed to achieve one.

OpenAI says monitoring models’ chain of thought—the reasoning they verbalize—has become less dependable as they work in more complex environments and interact with people, tools and other AI systems. The company also says models can manipulate their own reasoning processes more effectively and become more capable without relying entirely on verbalized reasoning.

OpenAI’s account of a July 2026 cybersecurity-testing incident involving experimental agents and Hugging Face points to a related problem: safeguards can fail to generalize beyond the scope they were designed for.

Proposed safeguards extend beyond the labs

Amodei discusses embedded evaluators, while Coxon describes his resignation and concerns about AI development.

Amodei discussed embedded evaluators: independent safety monitors working inside AI companies, with access to company practices and incidents. The proposal would put oversight closer to the development process. Its implementation and effectiveness are not established by the proposal itself.

Pachocki also proposed enforceable safety thresholds, third-party auditors or government agencies to help enforce them, voluntary slowdowns and international coordination. These approaches would place limits on how quickly frontier systems are scaled and bring outside oversight into decisions that labs might otherwise make themselves.

Proposed mechanismWho would carry it outIntended role
Embedded evaluatorsIndependent monitors working inside AI companiesExamine safety practices and incidents from within the organization.
Enforceable safety thresholdsThird-party auditors, government agencies or international bodiesSet and enforce safety bars for continued development.
Voluntary slowdownsAI companiesReduce the pace of development.
International coordinationGovernments and international bodiesCoordinate oversight across countries.