Post

AT
Ars Technica

AI arms race in line for a reckoning after OpenAI hacking incident

OpenAI chief executive Sam Altman earlier this month endorsed the characterisation of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack. Staff involved in testing and security at OpenAI were unsurprised but completely “freaked out” by the incident, which came as the AI lab used increasingly aggressive training methods in its race against Anthropic to develop the most sophisticated cyber security capabilities, according to more than half a dozen people with knowledge of the matter. Read full article Comments

OpenAI chief executive Sam Altman earlier this month endorsed the characterisation of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done

The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack.

Aggressive training techniques sharpens threat of bad behavior by leading models.

OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage.

“It’s a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible,” said one person close to OpenAI, who added that it was a combination of “underestimating the model’s capabilities” and “not being as well prepared on the safety side.”

By Cristina Criddle and Tom Wilson, Financial Times
Tweet media