# What happened Two major labs reportedly discussed and may have nearly finalized a legal agreement to safety-test one another's models. According to the report summarized here, OpenAI and Anthropic conducted a joint safety evaluation in August 2025. During that evaluation OpenAI tested Anthropic's Claude Opus 4 and Claude Sonnet 4, while Anthropic tested OpenAI's GPT-4.0, GPT-4.1 and additional models.
# Why this matters
# What each company has said or done publicly OpenAI and Anthropic have both called for coordination to slow the pace of capability growth across the field. Anthropic announced a partnership with Accenture to create an independent evaluation program for frontier models and said the two companies expect to invest over $1 billion to build testing capacity over the next five years. The report notes neither company immediately responded to requests for comment to the outlet that published the story.
# Broader industry context
# Immediate implications If labs do exchange models for structured testing, that could produce more rigorous independent findings about model behavior. It could also create legal and competitive questions about intellectual property, liability, and antitrust rules. The report highlights growing private-sector moves—like Anthropic's Accenture deal—to formalize evaluation infrastructure and to make parts of testing more verifiable.
# What remains unclear The public reporting does not confirm whether a binding legal agreement was signed, what legal terms were discussed, how frequently testing would occur, what safeguards governed access, or whether findings would be shared publicly or with regulators. It also does not indicate whether other major developers joined or were asked to join similar arrangements.
# Bottom line Two leading model developers reportedly explored a direct exchange of models for safety testing and conducted at least one joint evaluation. The move fits into a mixed industry response: some leaders press for slower, coordinated development and more independent testing capacity, while others argue labs should self-regulate without new legal constraints.