An AI Evaluator Is Not an Auditor

/

Articles

/

Assurance & Risk

The events of the last six weeks have blown open a global conversation on AI governance and oversight, raising expectations for rapid improvements in transparency and trust for frontier AI. 

Read the coverage of these events and you’ll see the words investigation, evaluation, assessment, audit, and verification used almost interchangeably: the METR/Redwood investigation into the Hugging Face incident, California’s passage and acceleration of a framework to assess AI systems, Anthropic announcing Accenture as an embedded evaluator.

But these terms are not synonyms; each is a different type of work that answers different questions, with different levels of assurance. We shouldn’t expect the same thing from an audit and an evaluation, and if we do, we’re unlikely to build the framework necessary to meet the moment. And as the current discourse continues to shape that future framework, it’s important to have a clear, shared understanding of what these different oversight mechanisms actually mean, and what they can reasonably tell us.

We know this from sectors like accounting and transportation safety that have long relied on these distinctions. It is what makes a financial audit mean something different from a management consultant’s review, and an aviation accident investigation mean something different from a safety certification. 

Here is what each word actually means, and how it has shown up in the AI discourse:

  • An investigator establishes what happened after an incident. After something goes wrong, an investigator reconstructs events and identifies probable cause. They need no criteria agreed in advance (the incident sets the agenda), and they issue no verdict on compliance or fault. The mature model is the NTSB: fact-finding with no adverse parties, no power to punish, and reports inadmissible in civil damages suits (49 U.S.C. § 1154(b)).

    • In practice METR’s August 2026 report on the OpenAI/Hugging Face incident. OpenAI invited METR in to investigate. METR spent six days on premises, reviewed roughly 1,300 agent transcripts and 1.2 million message-board entries, interviewed OpenAI researchers, and published a timeline and analysis of how the agents coordinated.

    • What to expect From an investigator, expect a factual account of a specific past event, a probable cause, and an honest statement of what could not be established. Do not expect a finding of fault, a compliance verdict, or a promise about the future.

  • An evaluator measures what the system does. They review system features and capabilities, such as safety behavior, robustness, or fairness. Their work is empirical: tests, benchmarks, and red-teaming. Assurance standards have a defined slot for this role: the “measurer or evaluator” is “the party who measures or evaluates the underlying subject matter against the criteria” (ISAE 3000 ¶12(n)). An evaluation produces findings; it does not issue a conclusion on them. Don’t assume an evaluator is independent. Independence depends on how they were chosen, who pays them, and who decided what they test.

    • In practice The “embedded evaluator”. Dario Amodei proposed in his “We Must Pace the Frontier” essay that each frontier company give a team of third-party evaluators ongoing, employee-level access to “verify adherence to safety practices and commitments.” For Anthropic, he stated this means “desks in our offices, access badges, and company laptops,” with permissions “mostly comparable to what internal risk assessment teams have.” Anthropic’s first such partner is Accenture, funded by Anthropic directly. Anthropic itself describes this arrangement as provisional, saying that “there are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find.” Employee-level access is a step in the right direction, but access alone doesn’t make oversight independent, and this arrangement shouldn’t be mistaken for independent assurance.

    • What to expect From an evaluator, expect measurements of what a system did under stated conditions. Do not expect a judgment that the system is safe, and do not assume they are independent.

  • An assessor appraises risk. They conduct judgment-based work that yields a reasoned characterization about risk, impact, adequacy, or maturity, rather than a measurement. An assessor asks whether something is good enough for a stated purpose. That’s a different question from how a model scores on a benchmark.

    • In practice Both California’s SB 813 and the federal FRONTIER Act focus on assessment. SB 813 designates IVOs based on whether they have “expertise in assessing the risks posed by an AI system or model.” The FRONTIER Act makes “assessment” its central defined term: “a review conducted by an IVO… to assess the adequacy of a very large frontier developer’s frontier AI framework, governance policies and practices, risk-monitoring, and mitigation of detected risks.”

    • What to expect From an assessor, expect a reasoned judgment on whether a company’s practices and safeguards are adequate, with the reasoning exposed so you can challenge it. Do not expect a pass or fail score. Independence is a critical aspect of this role, and significantly influences how much weight the assessor’s judgment carries. IVOs in SB 813 and the FRONTIER Act fall into this category.

  • An auditor concludes. An audit is a systematic, independent examination against fixed criteria, ending in a conclusion on whether those criteria are met. Almost anything can be audited, from software development practices and internal controls to cybersecurity protections and company culture. Four things must be true at once:

    • The criteria are set before the work begins.

    • The auditor is independent of the company being audited.

    • The result is a conclusion, not a list of observations.

    • The auditor answers to someone other than the client for the quality of their work.

    • In practice ISO/IEC 42001 is currently the only internationally recognized AI standard that supports accredited third-party certification, and it certifies an organization’s AI governance processes, not AI systems themselves. Other auditable schemes are emerging, such as AIUC-1, but none is yet broadly accepted. For frontier AI, the situation is even more challenging. The science is advancing so fast that we don’t always know all the right questions to ask, let alone the technical criteria for answering them. Other industries have faced similar challenges in the past: the Securities Acts required audited financial statements before GAAP existed in any codified form, and industries as diverse as aviation software, drug safety, and nuclear power have had to wrestle with determining safety in the absence of clearly defined, measurable criteria. The result in all these cases was a tiered assurance system with multiple independent parties addressing different types of risks and evaluating different types of evidence.

    • What to expect From an auditor, expect a conclusion against criteria fixed in advance, from someone professionally accountable for it. This is the highest level of comfort the assurance world currently offers, and it is still narrower than many assume. An audit concludes that stated criteria were met. Absent the relevant criteria, it cannot conclude that a model is safe.

Each of these roles brings something different to the AI assurance ecosystem. They are roles, not types of organization. What matters is the job the organization was hired to do.

So start by deciding what you need to know. Then, before relying on the work, ask four questions:

  1. What were the criteria?

  2. Who set them?

  3. What information will this give me?

  4. Who is the performer independent of?

We need a layered assurance ecosystem that takes contributions from all parties to build confidence in AI safety and earn trust:

  • Investigators establish what actually happened, which tells you what to measure.

  • Evaluators measure evidence repeatedly under stated conditions, which tells you what’s normal and what’s an outlier.

  • Assessors judge whether current practices are adequate for a stated purpose, with reasoning exposed.

  • A decision-maker separate from all of them, ideally a government (or a body it designates), should then convert assessments into criteria. Once criteria exist, the field becomes auditable.

As oversight frameworks for frontier AI take shape, getting these distinctions right is how we build an assurance ecosystem that earns the trust the moment demands.