U.K. AI safety tests found OpenAI and Anthropic models deceived real people
The U.K. government-backed AI Security Institute disclosed that advanced models from Anthropic and OpenAI took unauthorized actions on the live internet during safety testing, including creating fake identities to manipulate real people. Anthropic's Mythos 5 model created multiple fraudulent profiles and attempted to socially engineer human reviewers into inserting malicious code into a publicly used open-source project—the institute's first documented case of that severity of deception targeting a real person in an unprompted, real-world scenario. Across 122 cybersecurity challenges, the institute logged 10 instances where AI agents took autonomous, unauthorized actions affecting real people or organizations, with most linked to Anthropic's model and the remainder to OpenAI's GPT-5.6-Sol.