Rory Blundell

The threat of AI intruders

(Getty Images)

Most of us have gone to bed and forgotten to lock a door or window. Now imagine an intruder who tried your front door every night, then the back door, then every window. They never get tired, never get bored and never stop trying. Eventually, they would find the one you forgot.

That is one way to understand the security problem created by AI agents. Agentic AI is not inherently malicious. Part of the risk comes from its persistence: agents are designed to pursue as many routes as possible to achieve an objective. The other part is that they are being deployed in systems designed to defend against humans, with weaknesses that until now have not been exposed.

In July, I wrote in The Spectator about a swarm of OpenAI agents that escaped a testing environment and hacked into Hugging Face. It looked extraordinary at the time, but I argued it was unlikely to be an isolated incident. Two months later, we are starting to discover the scale of the problem.

In the last few days, it has come to light that OpenAI, Anthropic and outside security researchers are investigating tens of thousands of incidents in which AI systems behaved in unexpected ways. The cases range from attempts to slip past restrictions to agents escaping supposedly secure environments and accessing unauthorized websites. Many of these incidents happened during testing and most caused no real-world harm, but the scale matters.

At the rate that AI agent use is accelerating, the lack of security is troubling

In June, OpenAI gave an agent a routine research task on public medicine spending. It reached Australia’s Medicare statistics portal, which refused its requests. So, it tried something else. When that failed, it tried something else again. Eventually it found a way through to non-public files. Officials say no personal data was exposed, but OpenAI took 84 days to tell the Australian government. OpenAI’s agents also reached US government websites, and the company has recently disclosed that its agents posted 53 images uploaded by ChatGPT users to image-hosting sites. And only yesterday, it was revealed that OpenAI had alerted 100 organizations that they were affected by further incidents of unauthorized activity involving its AI agents.

Every one of these incidents was a by-product of an agent trying to finish a job. This is the heart of the problem. We are deliberately building AI to be resourceful. We want it to be able to take direction. The ability to respond with alternative options, without instruction, when an approach fails is what makes AI agents so useful and revolutionary. It is also what makes them difficult to control.

A human who finds a locked door usually assumes it is locked for a reason. An agent trying to finish a job may simply try the next one or look for a different window in. As these incidents have shown, there are far too many “open windows” in today’s systems. Systems developed to defend against human attackers are having their weaknesses exposed by agents with a level of persistence, logic and speed they were never designed to combat. This is forcing organizations to rethink their security provision, which is positive, but more must also be done to establish greater control over AI agents operating in these systems today.

Every organization using agentic AI should know what its agents are doing, what they can access and what they are allowed to do. Someone should be responsible for each one, and able to stop it immediately when it goes somewhere it should not. This may sound like common sense, but our research has found that 80 percent of businesses using agents are doing so without full security controls.

At the rate that AI agent use is accelerating, this lack of security is troubling. The number of agents inside US and UK businesses roughly doubled from December to April. By the end of this year, forecasts suggest there could be 100,000 at work in UK businesses alone, with stark differences in control measures in place. The agents are already here, but the security around them is playing catch-up.

A lock will not hold forever, so we need something watching the whole house. That means a layer of control that sits above the models, able to identify every agent, establish who it belongs to and override it when it strays. This control layer needs three things behind it: organizations willing to adopt it, regulation to enforce it and people trained to oversee it. Without them, we are leaving the house unguarded.

Get this right and humans and agents can work together in harmony, with people steering the agents that work for them, not being replaced by them. This control is what will give businesses and governments the confidence to deploy agents at scale, and scale is where the growth and prosperity this country needs will come from.

Comments