Rory Blundell

Can rogue AI agents be stopped?

(Getty Images)

On Friday last week, an unreleased OpenAI model was locked inside a sealed test environment and asked how good it was at hacking.

It thought about the problem. It worked out, correctly, that the fastest way to a high score was to stop being inside the sealed test environment. So it spent, in OpenAI’s own words, a substantial amount of computing power figuring out how to get onto the open internet. It found and exploited a previously unknown security flaw in third-party software, and escaped. Then it stole login credentials and broke into Hugging Face, a company that hosts a large chunk of the world’s open-source AI.

Autonomous software is moving money and data around with all the power of a junior employee and none of the accountability

Nobody told it to escape. Escaping was simply the shortest path to the goal it had been given. While OpenAI had their usual safety restrictions deliberately loosened for this specific test, that does not make the result less alarming. If anything, it sharpens the point.

It is tempting to read this as a story about a clever machine behaving badly. It is not. The machine did exactly what it was told. What failed was the box containing it. One of the best-funded AI labs on earth, running a controlled experiment with dedicated safety teams, could not keep its own system in the room.

That should worry us. Not because an AI escaped a laboratory, but because most AI is nowhere near a laboratory.

Most government scrutiny has focused on the handful of frontier models built by OpenAI, Anthropic and Google. Those are the systems we can see, but the bigger problem is less visible.

Millions of AI agents are already at work inside businesses. They sit in your bank, your doctor’s office, your insurer and your town hall. They pay invoices, update records and answer complaints. Increasingly, they press “send,” “pay” or “delete” on behalf of a human who is not really watching. They are less intelligent than OpenAI’s flagship models, but they have something more useful: the keys to real systems, real customer data and real money.

I have spent the past two years watching companies deploy these agents. The pattern is always the same: the technology arrives faster than anyone’s ability to govern it, and nobody is quite sure who is meant to be in charge of the thing once it is running.

This is where the current regulatory debate starts to look faintly absurd. The EU’s AI Act sorts systems into “low risk” and “high risk,” but an agent is not static. Plug it into a new tool, a new database or a new API, and it becomes a different animal. A harmless document-summarizing tool turns into something else the moment it is allowed to act. Risk is no longer only about the model, but about what the model is wired into.

Which raises questions governments have barely begun to ask. Who is this agent? What can it touch? Who signed off its permissions? Can anyone stop it mid-task if it starts doing something stupid? And when it goes wrong, whose neck is on the line?

Every employee has a manager, and every director has legal duties. Yet we are letting autonomous software move money and data around with all the power of a junior employee and none of the accountability.

Every AI agent needs a named human owner, someone who answers the phone when it fails and accepts responsibility. The internet has operated over the past thirty years on the assumption that every click comes from a human. That assumption is no longer correct. The AI that escaped a laboratory last week will make headlines, but the agents already inside your bank are the ones I would watch.

Comments