Anthropic's New 'Independent' AI Evaluator Looks a Lot Like Pre-Enron Auditing
Embedded evaluation beats an arm's-length audit on access. It's missing the four safeguards that make audits independent.
What Anthropic and Accenture actually announced
On September 18, Anthropic named Accenture as the first company to supply what it calls an “embedded evaluator”: a team with, as both companies describe it, employee-like access inside Anthropic. That means sitting in on training runs, following the decisions that shape how models get deployed, and talking directly to engineers, rather than reviewing finished output from outside. The work runs through Faculty, the AI-focused firm Accenture acquired earlier this year. Anthropic and Accenture are each putting at least $1 billion into the arrangement over five years.
The announcement is the first concrete move in a three-stage plan Anthropic CEO Dario Amodei calls “pacing the frontier,” laid out in a policy essay days before the Accenture deal. Reception was mixed; Nvidia’s Jensen Huang was reportedly among the skeptics. But the diagnosis behind the plan is one most people in AI safety already share: an outside audit, run on a schedule the audited company sets, with access the audited company grants, can only see what that company chooses to show. An evaluator working inside the building, with standing access, sees more.
That diagnosis is correct, and embedded access is a real improvement over an arm's-length audit. What's less settled is whether “embedded” and “independent AI evaluator” are the same thing, and the coverage of the deal — trade press and both companies' own newsrooms — spent almost no time on it. An evaluator needs more than a badge that lets it into the building. It needs a structure that makes it costly to stay quiet about what it finds.
Auditing already ran the "independent evaluator" experiment
Public company audits are built on a structure that should sound familiar: an outside firm, paid by the company it reviews, given access to internal records that outside investors never see. For most of the 20th century, American regulators treated that as sufficient independence, on the theory that an accounting firm's reputation was worth more than any single client's fees.
Arthur Andersen audited Enron for sixteen years. It was also one of Enron’s most highly paid consultants — in some years, Enron paid Andersen more for consulting than for the audit itself. When Enron collapsed in 2001, the reason Andersen had signed off on accounting that hid billions in losses was not a mystery: a firm that depends on a client for consulting revenue has a direct financial reason not to flag what its audit side finds. Access was never the problem. Andersen had plenty of it. The incentive structure wrapped around that access was the problem, and it took the largest bankruptcy in U.S. history at the time to get regulators to fix it.
The four fixes that came out of that failure
The Sarbanes-Oxley Act of 2002 did not ban company-paid auditors — doing so would have made public company audits nearly impossible to run. It kept the "audited company pays" structure and added four specific constraints most people never think about until they need to check whether a newer arrangement has them too.
One of those four is easy to overlook: the law also created a body to audit the auditors. Before 2002, an accounting firm answered to its own internal review process and, in theory, its professional reputation. The Public Company Accounting Oversight Board (PCAOB) gave outside regulators a standing right to inspect an audit firm’s work, independent of whether the audited company had any complaints. That detail matters here because "employee-like access" describes what the evaluator can see. It says nothing about who checks whether the evaluator is using that access well.
| Safeguard | Post-SOX financial audit | Anthropic-Accenture evaluator, as announced |
|---|---|---|
| Selling other services to the same client | Auditors are barred from selling most consulting services to companies they audit | Accenture's core business is AI strategy and implementation consulting; no wall with the evaluator team has been described |
| Who the findings answer to | The audit committee: independent directors, not company management | Not detailed publicly; described as working alongside Anthropic's own safety and governance functions |
| Public disclosure of findings | Audited financial statements and required disclosures are public by law | No disclosure commitment has been announced for evaluator findings |
| Oversight of the evaluator itself | The PCAOB inspects audit firms' work, not just the companies they audit | No independent body has been named to inspect the evaluators' own conduct |
None of this makes the deal hollow — it is days old, and companies often settle reporting lines after an announcement, not inside the press release. It means the four questions that decide whether “embedded” adds up to “independent” are still open, and they are the same four questions a different industry already had to answer under legal compulsion.
Access is not the same thing as independence
It's worth being precise about what embedded access actually buys. An evaluator who can watch a training run and talk to engineers directly will catch problems an outside auditor reviewing a model card never will. That's real, and worth the investment either company is making. What it doesn't do on its own is change who the evaluator answers to when what it finds is inconvenient. An evaluator who reports to the company, whose contract renews at the company's discretion, and whose findings the company decides whether to publish can be extremely well-informed and still have no reason to say anything the company doesn't want said.
“An evaluator with perfect access and no independent audience for its findings is a very well-informed employee, not a check.”
What would make this more than a press release
The version of embedded evaluation that would earn the word independent has a short, specific list: a public commitment to disclose material findings on a schedule the company doesn't control, a reporting line that runs to something other than the company's own management, a named boundary between Accenture's evaluator unit and its commercial AI consulting practice, and some form of outside inspection of the evaluators' own work — an AI-safety equivalent of the PCAOB. None of those four is exotic. All four already exist, tested under three decades of case law, in the profession that had to learn this lesson first.
Anthropic and Accenture have five years and roughly two billion committed dollars to fill in those details, and there's a real chance they do. “Pacing the frontier” is explicitly a three-stage plan; embedded evaluation is stage one, not the finished design. If the later stages add a public disclosure commitment, an independent reporting line, and some outside check on the evaluators themselves, this becomes the strongest AI safety mechanism any lab has built. If they don't, the word doing the most work in the announcement is “independent,” and it's the one word nothing in the deal currently backs up.
Whether the current announcement clears that bar isn't a question press coverage answers by repeating the word independent. It's a question that gets answered the first time an embedded evaluator finds something Anthropic doesn't want found, and the industry gets to see what happens next.
Frequently asked questions
Related reading
OpenAI Shut Down Atlas After 292 Days. Every Other AI Browser Is About to Learn Why.
OpenAI shut down its standalone Atlas browser on 9 August 2026, ten months after launch, even as usage was climbing. The real reason has more to do with Chrome’s grip on the desktop than with agentic AI.
Outcome-Based AI Pricing Didn't Fix the Metering Problem. It Just Put the Vendor in Charge of the Metric.
A seat is either logged in or it isn't — anyone can audit that. A 'conversation' or a 'resolution' is whatever the vendor's backend says it is. That's the part of the AI pricing shift nobody is pricing in.
The EU AI Act's High-Risk Rules Went Live on August 2. Article 12 Is the One Most AI Teams Will Fail.
August 2, 2026 made EU AI Act high-risk rules enforceable. Most guidance covers conformity assessments. The provision most AI teams will fail is Article 12, and it's an architecture problem, not a paperwork one.