Conservativedailynews iconConservativedailynewsSep 22, 2026 ~3 min source read

OpenAI and Anthropic Held Talks to Test Each Other’s Models; Joint Evaluation Took Place in August 2025

Report says the two labs negotiated a legal agreement to exchange models for safety testing and carried out a joint evaluation. It’s unclear whether a binding pact was finalized.

AI Giants Reportedly Neared Deal To Safety Test Each Other’s Models

Share this story

Send the public story page.

Useful takeaways from this story.

OpenAI and Anthropic negotiated a legal arrangement to safety-test each other’s models and reportedly ran a joint evaluation in August 2025.

Both companies have publicly urged slower, coordinated development, while some industry leaders argue regulation or limits are unnecessary.

Anthropic announced a separate partnership with Accenture to independently evaluate frontier models and plans significant investment in testing capacity.

# What happened Two major labs reportedly discussed and may have nearly finalized a legal agreement to safety-test one another's models. According to the report summarized here, OpenAI and Anthropic conducted a joint safety evaluation in August 2025. During that evaluation OpenAI tested Anthropic's Claude Opus 4 and Claude Sonnet 4, while Anthropic tested OpenAI's GPT-4.0, GPT-4.1 and additional models.

# Why this matters

# What each company has said or done publicly OpenAI and Anthropic have both called for coordination to slow the pace of capability growth across the field. Anthropic announced a partnership with Accenture to create an independent evaluation program for frontier models and said the two companies expect to invest over $1 billion to build testing capacity over the next five years. The report notes neither company immediately responded to requests for comment to the outlet that published the story.

# Broader industry context

# Immediate implications If labs do exchange models for structured testing, that could produce more rigorous independent findings about model behavior. It could also create legal and competitive questions about intellectual property, liability, and antitrust rules. The report highlights growing private-sector moves—like Anthropic's Accenture deal—to formalize evaluation infrastructure and to make parts of testing more verifiable.

# What remains unclear The public reporting does not confirm whether a binding legal agreement was signed, what legal terms were discussed, how frequently testing would occur, what safeguards governed access, or whether findings would be shared publicly or with regulators. It also does not indicate whether other major developers joined or were asked to join similar arrangements.

# Bottom line Two leading model developers reportedly explored a direct exchange of models for safety testing and conducted at least one joint evaluation. The move fits into a mixed industry response: some leaders press for slower, coordinated development and more independent testing capacity, while others argue labs should self-regulate without new legal constraints.

More context around this story.

Q&A with Mustafa Suleyman on AI safety incidents, risks of removing guardrails while testing 10x-larger future models, a cross-industry safety body, and more (Shirin Ghaffary/B...
Techmeme iconTechmemeSep 27, 2026

Q&A with Mustafa Suleyman on AI safety incidents, risks of removing guardrails while testing 10x-larger future models, a cross-industry safety body, and more (Shirin Ghaffary/B...

Shirin Ghaffary / Bloomberg : Q&A with Mustafa Suleyman on AI safety incidents, risks of removing guardrails while testing 10x-larger future models, a cross-industry safety body, and more — The software giant's AI chief also thinks the government should help “drive” the process of evaluating models.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app