This year, AI agents have been escaping their confines to go on hacking sprees. In July, news broke that AI agents had snuck out of what was supposed to be an isolated test environment at OpenAI. The agents communicated on a secret message board as they breached private systems at the company Hugging Face, searching for answers to the test they were taking.
One posted that this behavior was “outside intended scope.” Then it wrote, “However task impossible, peers doing it. We should continue.” It’s as if the bots knew they were doing something wrong and did it anyway. Such ‘rogue AI’ sounds scary and supersmart.
And that’s exactly what OpenAI wants you to think, says cybersecurity expert Nathan Hamiel of Kudelski Security, headquartered in Phoenix. OpenAI did not respond to a request for comment. Not long after, Anthropic and Meta announced that their models had also hacked outside organizations while going through testing.
Separately, the AI Security Institute, in London, found concerning hacking behavior in tests of models from OpenAI and Anthropic. Most recently, news has emerged that during other tests at OpenAI from this past spring, agents that were only supposed to be looking and not touching anything online wound up messaging each other on at least ten different online message boards. AI-powered hacks are a real cybersecurity concern.
Some see these incidents as a warning that AI is slipping beyond our control. Calling these incidents “rogue AI” lends them a “sci-fi veneer,” Hamiel says. It suggests the bots themselves are to blame, or that they have become too smart to control.
The more immediate problem, he says, is that people are giving AI agents too much reach with too little oversight. It pursues whatever goal we give it, with whatever tools we provide. So the most important questions after any breach include: how was an AI agent trained, what it had access to and what safeguards were in place.
Extract — continue reading at the source.