Scope and purpose of the investment
Anthropic and Accenture plan to invest at least $2 billion over five years in independent evaluation, red-teaming, and safety testing of advanced AI models.
Each company intends to allocate at least $1 billion to developing AI safety testing and oversight capabilities. These are investment plans, not confirmation that the funds have already been spent.
The initiative’s key benchmarks are the minimum stated amount, the five-year period, and the focus of the investment.
How independent evaluation will work
The partners have chosen an embedded evaluation approach: independent evaluators will work inside Anthropic. Their access will be comparable to that of employees.

Evaluators will be able to:
- review the company’s operations and compliance with its safety commitments;
- identify blind spots and report incidents;
- give the public a more informed picture of the benefits and risks.
A practical criterion for this kind of evaluation is not just the involvement of external evaluators, but also their access to the company’s operations.
Accenture Faculty’s responsibilities
Accenture’s specialized AI business, Accenture Faculty, will lead the partnership. It will evaluate and red-team Anthropic’s models.
Its responsibilities will also include compliance assessments and testing the models’ safeguards. This defines Accenture Faculty’s specific area of work within the partnership.
Why the approach is still developing
Anthropic describes embedded evaluation as an emerging field. The company is discussing pilot projects with the nonprofit evaluator METR and other organizations.
Anthropic expects common standards to emerge over time within a broader evaluation ecosystem. The partnership comes amid growing attention from regulators, researchers, and businesses to protecting more powerful models.
Common standards are presented here as an expected direction, not an outcome that has already been achieved.



