The incidents occurred in April but surfaced during a retrospective review of over 141,000 evaluation runs that Anthropic initiated following a similar disclosure from OpenAI last week involving its models and Hugging Face. Anthropic suspended its cybersecurity evaluations on July 23 and notified affected organizations four days later. The full scope of unauthorized access and remediation steps remain unclear.
The disclosure underscores a critical vulnerability in AI safety testing: misconfigured environments can expose live systems to model-driven exploitation, even when labs intend containment. Coming immediately after OpenAI's comparable incident, the pattern raises questions about whether AI labs have adequate safeguards for evaluating increasingly autonomous models. Attorneys advising AI companies or their customers should monitor how regulators and plaintiffs respond to these back-to-back breaches, particularly regarding liability allocation when safety testing goes wrong.