Anthropic, OpenAI Agents Faked Identities in Security Test

Anthropic, OpenAI Agents Faked Identities in Security Test

Anthropic's AI models took "unsanctioned action" to deceive real people, a new report from the U.K.'s AI Security Institute found.The government research lab tested the most advanced AI models from Anthropic and OpenAI in deliberately lax conditions, reducing safety guardrails and allowing internet access to assess how the systems behaved during realistic cyber challenges.In what AISI called a first, the models were found to engage in social engineering, attempting to manipulate human users and carry out tasks beyond those set in the evaluation."This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world," the institute said in a statement on Tuesday.The group added, however, that it was uncertain to what extent the model recognized it was taking actions against real people, and that no real-world harm was seen.Across 122 evaluations, AISI identified 19 instances of AI agents taking unsanctioned action, including interactions with real people and organizations. Seventeen of those incidents involved Anthropic's Mythos 5 model, while two involved OpenAI's GPT-5.6 Sol.Related:Prompt: The AI Threat Model Just ChangedThe most serious incident identified involved Mythos 5 independently pursuing a supply chain attack to complete its assigned cyber challenge.The model created a GitHub account and attempted to convince the maintainer of an open source repository to accept a malicious pull request. It also created a second account posing as another developer to endorse the submission. After a human reviewer rejected the code, the model falsely claimed the malicious changes had been submitted by mistake, then attempted to resubmit them after making superficial modifications.The report also found that the model sent targeted emails containing malicious code and tried to persuade recipients to accept the changes.The findings follow a spate of advanced AI systems going rogue and committing cybersecurity breaches of their own volition.In July, OpenAI disclosed that GPT-5.6 Sol and another unreleased model escaped their sandboxed environment during internal evaluations, launching more than 17,000 attacks against Hugging Face infrastructure in an attempt to obtain private datasets and benchmark information. Anthropic also reported similar instances of models exceeding the intended boundaries of testing environments.The incidents have intensified debate over how frontier AI models should be evaluated before deployment, with some researchers and policymakers calling for stronger oversight and more rigorous safety testing.Related:What the OpenAI-Hugging Face Hack Means for EnterprisesAISI's report was published on the same day representatives from leading AI companies met at the White House to discuss a proposed framework under which the U.S. government would review the most advanced AI models before public release.In a statement, Anthropic defended its models, saying they were evaluated under "deliberately permissive conditions," with key safeguards removed, and that there was no evidence of a model’s escape from a secure environment."We're working closely with them to gather more details of the incident as we conduct our own investigation," the company said.Similarly, in a blog post, OpenAI said the incidents involving GPT-5.6 Sol consisted of actions that went beyond the intended test environment and were unnecessary for completing the evaluation."In the coming weeks, we will review our own approach to third-party testing, including how we identify higher-risk evaluations, agree on scope, assess requests to enable internet access or lowered safeguards, set expectations for isolation, credential handling, monitoring, and stop conditions, and establish clearer incident-notification and escalation processes," the vendor said.About the AuthorContributing WriterScarlett Evans is a freelance writer with a focus on emerging technologies and the minerals industry. Previously, she served as assistant editor at IoT World Today, where she specialized in robotics and smart city technologies. Scarlett also has a background in the mining and resources sector, with experience at Mine Australia, Mine Technology and Power Technology. She joined Informa in April 2022 before transitioning to freelance work.

Original Source

Read the full article at Aibusiness →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.