AI arms race in line for a reckoning after OpenAI hacking incident
OpenAI chief executive Sam Altman earlier this month endorsed the characterisation of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack. Staff involved in testing and security at OpenAI were unsurprised but completely “freaked out” by the incident, which came as the AI lab used increasingly aggressive training methods in its race against Anthropic to develop the most sophisticated cyber security capabilities, according to more than half a dozen people with knowledge of the matter. Read full article Comments
OpenAI chief executive Sam Altman earlier this month endorsed the characterisation of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done
The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack.
Aggressive training techniques sharpens threat of bad behavior by leading models.
OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage.
“It’s a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible,” said one person close to OpenAI, who added that it was a combination of “underestimating the model’s capabilities” and “not being as well prepared on the safety side.”