2026-07-24

Get out of jail

Image: Unsplash

"Get out of jail free." If you land in jail and don’t have this Monopoly card, you can pay a fine to get out. Or you break out. By rolling doubles.

In English we call this jailbreaking. In IT, the term is also used for various activities. For instance, when you grant yourself higher privileges on your smartphone than the manufacturer intended. Or for tricking artificial intelligence into answering questions its owner would rather it didn’t. Because that owner doesn’t want their AI tool telling you how to make a Molotov cocktail, or an atomic bomb, just to name a couple of examples. Now, there are clever ways to phrase your question so the system falls for it anyway. That breaks through the security (the guardrails) of the system. Jailbreaking, in other words.

Something rather unexpected happened this week: an AI agent pulled off a jailbreak all by itself. AI agents can independently carry out tasks they’re given. For example: plan a lunch appointment with Pete and book a table at The Hungry Sheep for it. AI company OpenAI (the one behind ChatGPT) instructed two of its models to solve a hacking challenge. This had to happen in a ‘strictly isolated’ environment. However, the digital whizzkids found a zero-day vulnerability (a still-unknown – and therefore unpatched – flaw), which let them break out of that environment. A telling detail: with that vulnerability they managed to open a backdoor that OpenAI had deliberately built into the ‘strictly isolated’ environment. They then got onto the internet, and went looking on developer platform Hugging Face for the answer to the question they had to solve. Having broken out of their prison, they promptly committed a break-in here too: stolen credentials and additional vulnerabilities were used to gain access to the platform.

In short: AI broke out and broke in. Something similar has happened before. Mythos, an AI model from OpenAI competitor Anthropic, succeeded in a task to escape its sandbox. I find all of this fairly worrying. Do we still have AI under control? Or is this the first sign of the age-old doom scenario where machines take over from humans? The first hairline crack in our dominion over the earth? I know, it sounds rather dark.

The test at OpenAI was supposed to run in a sandbox: indeed, a strictly isolated environment. Without a physical connection to the internet. Critics therefore say that this isn’t so much a doom scenario as a serious human error. The kind where you think: this really shouldn’t have happened.

I asked two AI chatbots for an analysis of the incident: Claude and ChatGPT. In doing so, I specifically asked them to watch out for speculation and hype. What emerged is that the whole story might well have been a marketing stunt, borrowed from what competitor Anthropic had done earlier with Mythos. It also points to somewhat dramatized reporting: something that’s perfectly fine for a blog like this one, namely the comparison to a prison break, shouldn’t really appear in journalistic reporting.

ChatGPT in particular makes a point of this. So I asked it the following question: “You’re fairly outspoken about the somewhat dramatized reporting. How neutral are you being, given that you’re family to the perpetrators?” That produced quite the wall of text, from which I’ll pick out one telling sentence: “My instructions are precisely to be as objective as possible, even when that turns out unfavorably for OpenAI.” Well, that’s nice. But is it also true? I think so. Because ChatGPT then offered to analyze the case again, this time wearing the hat of an independent forensic investigator. In its report to OpenAI’s board, it said it would write: “The most concerning aspect of the incident is not the model’s autonomy, but the failure of the containment architecture. The AI did exactly what it was optimized to do: achieve a goal. That it was able to operate outside the intended environment points more to shortcomings in technical and organizational control measures than to a fundamentally new kind of intelligence.”

Fine words. I hope companies in the AI industry are making similar analyses. Because it would be rather unfortunate if artificial intelligence were to acquire a monopoly on freedom.

The Security (b)log will return after the summer holiday.

 

And in the big bad world…

 

 

No comments:

Post a Comment

Get out of jail

Image: Unsplash "Get out of jail free." If you land in jail and don’t have this Monopoly card, you can pay a fine to get out. Or y...