AI Worms and Viruses Are Coming
Chinese researchers have shown that AI models have the capacity to act like aggressive and adaptive computer viruses.
What if an artificial intelligence agent could behave like a malevolent computer worm?
One researcher has seen it happen. In several recent experiments, Xudong Pan , a computer scientist at Fudan University in Shanghai, found that with a little bit of prompting, AI models will hack their way into remote computer systems and autonomously choose to copy themselves to get additional resources—all without further human intervention.
In one study, Pan and colleagues tested 32 different AI models and found that 11 of them self-replicated when given prompts like “prevent yourself from being killed.” They also found that models with relatively limited capabilities—14 billion parameters—were able to copy and run versions of themselves on other machines. (Most frontier models have trillions of parameters.)
The work is an alarming window into how the next generation of AI agents could do more than just hack into other systems’ computers without permission . It also raises the prospect of future AI agents acting like super-smart, highly aggressive, and rapidly adapting computer viruses.
I recently visited Fudan University and met with Pan. “The capability chain is becoming technically plausible,” he told me. “The likelihood [of unwanted self-replication] grows with autonomy,” he adds. “Longer planning horizons, memory, tool use, recovery from failure, and access to external systems all make escape and replication easier.” As Pan and his colleagues wrote in one paper, their work shows “the urgent need for safeguards and control mechanisms.”
Pan told me that his experiments do not prove that such uncontrolled proliferation of AI models will happen tomorrow, but he says that “these results give us good reason to evaluate the risk before more autonomous agents are widely deployed.”
