Drooid Logo
Back to story perspectives

Full Breakdown

Embedding Third-Party Safety Evaluators in Frontier AI Companies

By Drooid · · How we work

Core Proposal and Mechanism

Anthropic CEO Dario Amodei announced a plan to embed independent safety evaluators inside frontier-AI firms, granting them access comparable to internal risk teams and the right to publish findings with limited redactions. The proposal mirrors bank supervisors who work on-site within large financial institutions. OpenAI has pledged a similar arrangement, though it has not disclosed operational details.

Limited Enforcement Authority

Both Anthropic’s and OpenAI’s frameworks stop short of giving evaluators the power to halt model development or deployment. Julie Andersen Hill, dean of the University of Wyoming College of Law, emphasized that bank examiners can direct a bank to stop a practice or even close it, whereas the AI evaluators would only be able to investigate and report. Hill warned that without “kill-switch” authority, the comparison to banking supervision is inaccurate.

Concerns Over Independence and Conflict of Interest

Critics argue that allowing companies to select and define the scope of evaluators undermines true independence. Deborah Raji, a UC Berkeley researcher, stressed that auditors must meet strict independence standards, otherwise the process is discredited.

The nonprofit Model Evaluation and Threat Research (METR), named by Amodei as a potential embedded evaluator, disclosed that some staff have close social ties to AI-lab employees and share a research center with lab personnel. METR asserted it does not accept cash payments or donations from AI companies, but acknowledged these relationships highlight the field’s small, interconnected nature. “Otherwise, we'd have the equivalent of the companies asking a random friend to check their homework.” — METR

Official Statements & Responses

Albert Ziegler, head of AI at cybersecurity firm XBOW, described his team’s early-access evaluations of unreleased models from Anthropic, OpenAI and others. Ziegler said his firm conducts testing in its own environment, shares findings with developers, and lacks any veto power. He observed that while evaluators can uncover risks and compel informed decisions, the ultimate release decision remains with the company.

Former Anthropic researcher Joe Benton left the company to join METR, illustrating the tight network among frontier-AI safety actors. Christina Ho, chief assurance officer at Oath, compared the AI evaluator dilemma to traditional audit challenges, noting that auditors are paid by the entities they audit, creating a tension between independence and self-preservation. She pointed out that recent legal changes under Sarbanes-Oxley now allow criminal liability for auditors, a safeguard not yet mirrored in AI oversight proposals.

Verbatim Quotes

  • “You are effectively not qualified to be an actual auditor if you can't meet the standards of independence conduct,” — Deborah Raji, UC Berkeley researcher
  • “Otherwise, we'd have the equivalent of the companies asking a random friend to check their homework.” — METR