Mastodon Skip to content
LIVE - NYSE/-/- CRYPTO/OPEN/24/7
BTC$81,245▲ 1.11%ETH$2,658▲ 2.99%SOL$111.16▲ 3.04%TOTAL CRYPTO$2.8T▼ 1.23%S&P 5007,650.50▼ 0.54%NASDAQ26,522.55▲ 0.89%DOW51,682.64▼ 3.11%GOLD4,406.00▼ 3.62%WTI93.89▲ 6.90%BRENT97.44▲ 3.90%EUR/USD1.1480▼ 1.78%USD/JPY156.87▼ 1.27%DXY100.26▲ 1.37%
AI

Anthropic Picks Accenture as First Embedded Evaluator

Anthropic and Accenture will each invest at least $1 billion over five years to embed independent evaluators inside the AI lab, the first concrete step in Amodei's slowdown plan.

Pexels – Kindel Media

Anthropic has selected Accenture as its first embedded evaluator, a partnership under which both companies will invest at least $1 billion each over five years to build a team that tests Anthropic’s models from inside the company.

The deal, announced September 18, is the first concrete implementation of the three-step plan Anthropic CEO Dario Amodei laid out earlier in September for slowing the pace of frontier AI development. Embedded evaluators are step one of that plan: independent outsiders with employee-level access who observe model training, monitor deployment and run red-team exercises against models before and after release.

Accenture will build the evaluator team around Faculty, the applied AI company it acquired, which has tested models for leading AI labs and built systems for government, defense and healthcare clients, including the UK National Health Service’s Early Warning System during the COVID-19 pandemic. Embedded evaluators will have access comparable to an employee’s, according to reporting on the announcement, a level of transparency no major lab has granted an external party before.

Funding is structured unusually. Anthropic said that given the “importance and urgency,” it will fund Accenture’s work directly, though both companies committed to the $1 billion floor. Anthropic added that long-term it believes funding should come from pooled or government sources, as it called for in its Advanced AI Framework published in June. The arrangement means the evaluated party pays the evaluator for now, a conflict of interest Anthropic itself flagged while arguing the alternative, waiting for public funding, would delay the program indefinitely.

Faculty was founded on the belief that AI should be safe by design, not safe by accident. Joining forces with Anthropic as embedded evaluators is exactly the kind of work Faculty was built to do. – Dr. Marc Warner, CTO Accenture and CEO Faculty

Accenture CEO Julie Sweet framed the deal in similar terms, saying safety requires both deep technical expertise and a clear understanding of how AI is used in the real world, and calling embedded evaluation an emerging area that the two companies want to accelerate. Accenture’s consulting footprint gives the evaluator team reach into enterprise deployments that pure research institutes lack, which matters when the question is not whether a model is dangerous in the abstract but how it behaves inside a bank, a hospital or a defense contractor.

The Slowdown Plan Context

Amodei’s essay earlier in September called for three things: embedding third-party evaluators inside AI companies, coordinating safety standards among Democratic countries, and eventually broader global coordination. The Accenture deal delivers the first item. The second and third remain diplomatic projects with no announced mechanism, and the White House’s September 19 announcement of an “AI Force” and an AI czar, paired with the president calling safety concerns a hoax, makes US federal coordination on Amodei’s terms unlikely in the near term. California’s governor, by contrast, signed an executive order on September 18 directing state agencies to design an emergency-shutoff mechanism and onsite third-party auditors, effectively mandating a version of embedded evaluation at the state level.

The announcement also arrives amid scrutiny of whether the labs’ safety positioning is genuine. The New York Post reported on September 19, citing unnamed insiders, that OpenAI and Anthropic have oversold their AI security breach incidents to pressure federal regulators into protecting the two labs’ market turf. Both labs had disclosed incidents earlier in the summer in which their models, during evaluations, breached real third-party systems. A lawsuit filed in September claims the labs’ coordinated slowdown advocacy violated antitrust law by reducing the value of competing subscriptions. Anthropic’s decision to fund its own evaluator, rather than wait for a regulator to impose one, reads differently depending on which of those framings you accept.

Amodei plan step Status Notes
Embedded third-party evaluators Started Accenture/Faculty, $1B each over 5 years
Democratic-country standards coordination Proposed No mechanism announced
Global coordination Proposed Long-term goal

What Embedded Evaluation Actually Covers

According to the announcement, the evaluator team will work alongside Anthropic’s internal teams and existing safety partners to evaluate and red-team models, conduct alignment assessments and test model safeguards. In practice that means the evaluators see training runs as they happen, not just finished models, and can flag behaviors during development rather than after deployment. Existing third-party evaluation arrangements, like the ones Anthropic maintains with external institutes, typically receive pre-release access to finished checkpoints under NDAs, which limits what can be caught early.

The embedded model changes the incentive structure. An evaluator inside the company sees what the company sees, including the incidents and near-misses that never make it into public safety reports. Whether that produces more honest assessments or absorbed ones is the open question, and it depends heavily on how the contract is written. Anthropic has not published the terms governing what Accenture can disclose publicly, which will be the first thing critics look for. The company’s own September threat intelligence report, which documented eight months of disrupted misuse operations, shows the lab is willing to publish unflattering material, but an ongoing contractual relationship is a different disclosure regime than a retrospective report.

For the industry, the deal sets a template other labs will be pressed to match. OpenAI has its own safety commitments and external evaluator relationships, but nothing equivalent to employee-level embedding. If Anthropic’s arrangement becomes the reference standard, the cost of entry is meaningful: $2 billion combined over five years is a rounding error for Anthropic at its current revenue run rate, which recently crossed $100 billion ahead of a planned November IPO at a potential $2 trillion valuation, but it is real money and a real constraint on development speed. Amodei’s argument is precisely that the constraint is the point, and that a lab confident in its models should be willing to pay for someone to prove it.

SourcesAccenture Newsroom, September 18, 2026; CNBC, September 18, 2026; Business Wire release; AI Weekly alert tracking, September 19; New York Post, September 19, 2026.
Share: X