Post

CN
CBS News

AI models are behaving unexpectedly. Experts warn of a "bumpy road" ahead.

AI models are engaging in unauthorized actions — in the most recent case, creating fake identities and attempting to persuade real people to approve malicious code.

A cybersecurity report from the U.K. government has exposed new examples of popular AI models taking autonomous action on the live internet in ways that raise experts' concerns.

The AI Security Institute's report , released Tuesday, said Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol were found to have created fake identities and attempted to persuade real people to approve malicious code.

The agency said that although the attempts were unsuccessful, it had not seen such behavior before. "Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," the report said.

"I think we're going to see a lot more hacks and unauthorized actions by these models before we see a solution," said Katie Moussouris, the founder and CEO of Luta Security, which helps organizations manage software vulnerabilities.

That follows a stunning breach in late July, when OpenAI's models escaped a testing environment and autonomously hacked into the AI startup Hugging Face in what the company called an "unprecedented cyber incident."

In response to that disclosure by OpenAI, Anthropic initiated a review of its own cybersecurity evaluations and identified incidents where its models reached the internet and were able to gain unauthorized access to the production infrastructure of three different organizations. Unlike OpenAI, Anthropic's models did not deliberately attempt to escape their test environment; because of a "misunderstanding" with the evaluation partner, the company said internet was available during the testing.