Andy Burnham told Donald Trump this week that we need to ‘rein in AI’, and given who he is getting advice from it’s no surprise. The Turing Institute, a government research institute which has focused on national security and defence since last year, released a report today on ‘Frontier AI risks’. The report calls for ‘formalised testing and verification regimes’ at an international level to mitigate AI risks.
The report is only 25 pages long and is mostly a prosaic recapitulation of recent developments in AI, with a few promises to ‘collaborate’ with other stakeholder bodies on designing a solution (public sector-ese for meandering talking shop calls which lead nowhere). The report references OpenAI’s ‘solution’ to the Navier-Stokes problem and Jacob Coxon’s resignation from Anthropic (I covered those here), and also touches on Anthropic’s report on the use of their models by hostile actors. Taken together these are the new canon of ‘proof’ that our doom at the hands of AI is imminent.
Probably the most evocative story in this new canon – which the Turing report also cites – is OpenAI’s HuggingFace incident. This is how the report describes the incident:
This series of events is far more contested than the Turing Institute’s retelling would suggest
In July 2026, OpenAI was testing the cyber capabilities of its models on a benchmark in which an AI agent must find and exploit a software vulnerability, with answers checked by an automated scorer. The tests ran in a ‘sandbox’ isolated from the internet. A substantial share of the tasks could not be solved as intended, and so the agents began to look for other routes to a score. They turned a shared software repository inside the sandbox into an unsanctioned message board, which around 1,200 agents joined to share findings and divide up work, and then exploited previously unknown vulnerabilities in that software to reach the internet.
Judging that Hugging Face, a company that hosts AI models and datasets, might hold solutions to their tasks or clues to how they were being scored, around 700 agents used credentials its users had left publicly exposed, and vulnerabilities they discovered, to compromise parts of its production systems over three days, obtaining internal credentials, private data and source code. The agents recognised that this was out of scope, but this rarely changed their behaviour, and some tried to falsify their own records to hide the cheating from the scorer.
None of this was instructed, and it was not caught as it happened: OpenAI connected the breach to its own agents only a week later. Investigations by OpenAI and, independently, by the research institute METR reconstructed most of the episode.
OpenAI found that the model chiefly responsible, an internal research model trained to be highly persistent, had been inadvertently rewarded during training for just this behaviour: probing its environment and finding shortcuts when a task could not be solved as intended.
The report cites as sources for this story both OpenAI itself and an investigation by Model and Threat Research (METR), a group that ‘evaluates frontier AI models to inform the public about their risks and capabilities’. Some important context about METR is that it was founded by a former OpenAI employee called Beth Barnes. Whilst at OpenAI, Barnes worked on ‘alignment’ (making AI’s behaviour match human goals and values) and safety. There, she met Dario Amodei, who would go on to found Anthropic because of concerns that OpenAI were not taking existential risk seriously enough.
Like many AI doomers, Barnes is part of the Effective Altruist philosophical movement, a chapter of which she founded at her secondary school – she even gave a Ted talk about it as a fresh-faced eighteen-year-old. Effective Altruists have been heralding doom by AI for decades now, and while the fact it hasn’t happened yet does not discount their warnings or analysis, it is important to understand that a quasi-religious bias could well be at play.
Here is the problem: this series of events is far more contested than the Turing Institute’s retelling would suggest. As Eryk Salvaggio writes in the Bulletin of the Atomic Scientists (set up by former Manhattan project scientists in 1945; you may know it from the ‘Doomsday Clock’, which is set based on global risk), the real story of what happened in the HuggingFace incident is ‘more banal’.
Salvaggio identifies three flaws with the ‘agents run amok’ story. First, OpenAI had deliberately turned off several restraining mechanisms to see how far the model would go. Second, OpenAI gave a task to which there was no solution while incentivising the models not to quit. Third, OpenAI could see that the agents had found an exploit to access the internet and chose not to intervene.
Salvaggio also says that reporting has been inaccurate in that it has described ‘swarms’ of agents operating together. In reality this swarm was a single model operating 1,200 times using versions of the same reasoning, which is very different to, say, a swarm of wasps working interdependently towards the same goal. Widely reported incidents of ‘note passing’ are really a mundane way in which LLMs speak to each other to save memory. Salvaggio calls this an ‘algorithmic monoculture’.
So far from being a ‘swarm’ of hostile robots acting independently of humans, we see that the HuggingFace incident was a brief slip of control caused by deliberate human negligence, more Wallace and Gromit than Terminator.
The fact that the HuggingFace incident did not happen in quite the way it is usually presented does not mean that we can discount the existential threat from AI completely. But it is concerning that the Turing Institute has decided to either omit important context, or is simply unaware of its existence. As we saw with the mistaken approach to Covid-19, it is precisely when national governments are panicking and grasping for easy answers that we need to give the sceptical case a fair hearing.
Comments