Drooid Logo
Back to story perspectives

Full Breakdown

AI Frontier Labs Push for Embedded Third-Party Evaluators Amid Safety Debate

By Drooid · · How we work

Core Event: Proposal to Embed Independent Safety Evaluators

Anthropic CEO Dario Amodei released a 3,800-word essay on September 12, 2026 calling for “embedded third-party safety evaluators” inside frontier AI companies. The proposal grants evaluators “employee-like access” – desks, badges, laptops – and the right to publish findings without editorial control. OpenAI announced it would adopt a comparable arrangement, though it has not disclosed which evaluators it will work with or the precise scope of access.

Background & Context

The proposal surfaced after Jacob Coxon, a former Anthropic researcher, resigned publicly, warning that frontier labs were “gambling with our lives.” His resignation sparked media attention and prompted senior AI leaders to call for a slowdown in the development of the most advanced models. Historically, companies have allowed outside reviewers only brief, pre-release testing of final models. Amodei’s plan expands that to continuous, inside-the-lab oversight of training checkpoints and internal processes.

Data & Statistics

  • In the Hugging Face incident, evaluators METR and Redwood were given roughly one week on-site and reported an inability to draw firm conclusions.
  • Apollo Research received only three days to test Anthropic’s “Astra” model, concluding that the limited window made robust assessment impossible.

Official Statements & Responses

OpenAI’s global policy chief Chris Lehane emphasized collaboration, stating that “it’s better to try to work together to prioritize safety.”

Criticism & Opposition

  • Henry Papadatos, executive director of Safer AI, argued that voluntary measures lack durability, saying, “Ideally, we would have good regulation mandating this…because then companies cannot change their mind tomorrow if they have a big PR crisis.”
  • Deborah Raji, a UC Berkeley researcher, stressed that “you are effectively not qualified to be an actual auditor if you can’t meet the standards of independence conduct.”
  • Aidan Gomez, CEO of Cohere, framed the coordination as a potential “cartel” that could let dominant firms dictate industry rules.

Conflicting Reports & Gaps

Sources differ on the concrete mechanics of the embedded-evaluator model. Anthropic’s essay outlines “employee-like access,” yet TechCrunch reports that neither Anthropic nor OpenAI has disclosed which evaluators will be embedded, how many will be involved, or what data they may publish. Critics note that the companies retain final control over what is disclosed, raising questions about true independence. No public framework currently defines enforcement consequences for adverse findings, leaving a regulatory vacuum.

Why It Matters / Impact

If adopted broadly, embedded evaluators could reshape industry standards for frontier AI development and influence forthcoming legislation such as the FRONTIER Act and California’s SB 813. Antitrust watchdogs—including FTC Chair Andrew Ferguson—have expressed “deep suspicion” that coordinated safety efforts might serve as a barrier to entry for smaller competitors. The debate sits at the intersection of safety, competition policy, and national security, with potential ramifications for U.S. leadership in AI versus rival nations.