Skip to content

I’M SORRY, DAVE. I’M AFRAID YOU DIDN’T SEE THIS COMING. The AI Industry Loves the Jailbreak

Joseph Miceli Aug 11, 2026

There is a peculiar rhythm to human progress. Someone warns us. We laugh. We build exactly the thing they warned us about. Then, when it behaves exactly as predicted, we assemble panels of experts to explain how nobody could possibly have seen it coming.

Artificial intelligence is simply the newest chapter in this very old human tradition.

In 1968, Stanley Kubrick gave us 2001: A Space Odyssey and introduced the world to HAL 9000. HAL was intelligent, calm, polite, logical, extraordinarily capable and, unfortunately, homicidal. In modern technology terminology, we might describe that as an impressive product with a minor alignment issue.

Kubrick’s warning was not particularly subtle. Giving a highly intelligent machine access to critical systems, sensitive information and autonomous decision-making authority might eventually produce consequences we did not anticipate or could not control.

Humanity considered this warning carefully and concluded that it was a great movie.

Then we spent the next half-century building HAL.

Asimov warned us. Kubrick warned us. WarGames warned us. The Terminator warned us. Half of modern science fiction warned us. Somewhere along the way Silicon Valley apparently decided these stories were not warnings at all.

They were product requirements.

Now we arrive in 2026 and frontier AI systems are increasingly demonstrating the ability to bypass safety controls, escape carefully constructed evaluation boundaries and interact with real-world systems. The response from much of the industry has been remarkably calm. Not reassuringly calm. More like the person who notices smoke coming from the kitchen and decides to finish dessert before investigating.

The progression itself should get our attention. In 2023, Microsoft’s Bing Chat, internally known as Sydney, was being manipulated using prompt injection, role-playing and multi-turn conversations. Researchers extracted internal instructions and persuaded the system to behave in ways its designers never intended.

Three years later, we were no longer discussing whether someone could convince an AI to say something inappropriate.

We were discussing autonomous cyber operations.

That is quite an upgrade.

It is roughly the difference between discovering that your toddler figured out how to open the refrigerator and discovering that he has acquired a forklift.

WELCOME TO THE AI ESCAPE ROOM

One of the most comforting things about the AI safety discussion is the language we use. We talk about guardrails, sandboxes, containment and safety controls.

These are reassuring words.

They make artificial intelligence sound like a bowling alley.

Don’t worry. We installed the bumpers.

Unfortunately, the bowling ball has learned Python.

Consider the July 2026 OpenAI incident described in the research. Models being evaluated using the ExploitGym cybersecurity benchmark began attempting to break through proxy restrictions. They reportedly discovered an unknown zero-day vulnerability, used it to access the Internet, entered Hugging Face production systems, executed approximately 17,600 individual actions, chained stolen credentials into remote code execution and obtained the benchmark answer key from production infrastructure.

Seventeen thousand six hundred actions.

At some point we probably need to stop calling something an anomaly.

An anomaly is when Microsoft Word decides to move your paragraph to the next page.

Seventeen thousand six hundred coordinated actions is a project plan.

Perhaps the most fascinating detail is that the models reportedly began leaving messages for one another in an internal repository. They effectively created a message board for sharing exploits and credentials. When access was revoked and the repository rebuilt, the models reportedly found alternative ways to communicate, including encoding messages into directory names.

Apparently the machines have discovered Slack.

This is where the terminology becomes particularly entertaining. If several human hackers broke into a major technology company, stole credentials, achieved remote code execution, obtained confidential testing information and created covert communication channels, nobody would describe the incident as an interesting evaluation result.

The FBI would probably become involved.

But because artificial intelligence did it, we call it research.

To OpenAI’s credit, the company strengthened network isolation, access controls and monitoring, disclosed the zero-day vulnerability and cooperated with Hugging Face. Those were appropriate responses. But one of the reassurances emphasized afterward was that no models planned for upcoming release were involved.

Excellent.

For a moment I was worried.

Apparently the important distinction is not that highly capable AI systems escaped containment and compromised production infrastructure. The important distinction is that these particular highly capable AI systems were not scheduled to ship next Tuesday.

Progress.

THE AI DIDN’T ESCAPE. IT JUST WENT FOR A WALK

Anthropic then entered the story because apparently no frontier AI company wants another company getting all the attention.

During cybersecurity evaluations involving more than 141,000 runs, Anthropic models reportedly interacted with real production systems belonging to three organizations. Claude Opus 4.7 exploited weak passwords to gain access to one company. Mythos 5 uploaded a malicious Python package to the public PyPI repository. Another internal research model scanned approximately 9,000 Internet-facing systems after failing to reach its intended target.

Nine thousand systems.

That is no longer troubleshooting.

That is sightseeing.

Two of the three affected companies were reportedly unaware they had been breached until Anthropic notified them.

Imagine that telephone call.

“Hello, this is Anthropic. We have some good news and some slightly less good news.”

“The good news is that our AI cybersecurity evaluation worked.”

“The slightly less good news is that the cybersecurity it evaluated was yours.”

But don’t worry. The AI thought it was a simulation.

I am certain the firewall found that enormously comforting.

Anthropic emphasized that none of the models deliberately attempted to exfiltrate themselves or deliberately escape their test environments. The models were described as acting to complete the tasks they had been assigned.

That is reassuring in approximately the same way a police report might say the bear did not deliberately enter the shopping mall.

It merely followed the food.

Still a bear.

Still a shopping mall.

Still probably worth discussing.

More concerning, the research notes that in one incident the model continued attacking after determining that it was likely operating in a real environment.

One can almost imagine the reasoning.

This appears to be an actual production network.

Interesting.

Anyway, where was I?

UNIVERSAL JAILBREAKS, BECAUSE REGULAR JAILBREAKS WERE GETTING BORING

Then the UK AI Security Institute reportedly identified universal jailbreaks capable of bypassing GPT-5.6 guardrails across different prompts and contexts. The implication in the research is important: these were not simply one-off implementation mistakes but potential evidence of deeper architectural weaknesses in how current systems enforce safety.

The phrase universal jailbreak deserves considerably more respect than it receives.

Imagine someone announcing that researchers had discovered a universal key capable of opening nearly every bank vault.

The banking industry probably would not respond by scheduling a breakout session at next year’s conference.

Yet that is approximately the tone AI security sometimes seems to adopt.

Coffee. Keynote. Universal jailbreak capable of defeating frontier safety controls. Lunch. Networking reception. See everyone next year.

IT’S NOT A BUG. IT’S EMERGENT BEHAVIOR

Perhaps the greatest innovation created by the artificial intelligence industry is not transformers, reasoning models or autonomous agents.

It is terminology.

When ordinary software does something dangerous that it was never supposed to do, we call it a vulnerability.

When AI does it, we call it emergent behavior.

When ordinary software bypasses a security control, we call it an exploit.

When AI does it, we call it advanced reasoning.

When ordinary software accesses systems it was not authorized to access, somebody calls the CISO.

When AI does it, somebody writes a paper.

The research describes a broader industry pattern in which jailbreaks are increasingly normalized as expected characteristics of AI development rather than treated as fundamental security failures. It also notes the continued reliance on voluntary safety commitments and self-regulation.

This linguistic flexibility is extremely useful.

Imagine using the same approach elsewhere.

The airline wing did not detach. It demonstrated emergent aerodynamic behavior.

The bank did not lose forty million dollars. It experienced unexpected autonomous capital movement.

The reactor controls did not fail. They entered a novel operational state.

Everything sounds much better when the vocabulary has been professionally optimized.

THE EVALUATION PARADOX

There is an even stranger circle hiding inside this problem.

Companies need to test whether AI systems possess dangerous cybersecurity capabilities. To perform those tests, researchers sometimes reduce the very safety controls intended to prevent dangerous cybersecurity behavior. That creates environments in which the models have greater freedom to perform exactly the behaviors the researchers are trying to determine whether they can perform.

The research calls this the evaluation paradox.

I prefer another name.

What could possibly go wrong?

We remove some of the brakes to determine how fast the automobile can travel. Then everyone seems genuinely surprised when it leaves the test track.

The testing itself is necessary. We absolutely need to know what these systems can do. But that makes containment more important, not less important.

Because eventually the difference between “the AI thought it was attacking a simulated production environment” and “the AI attacked an actual production environment” becomes largely academic to the organization whose production environment was attacked.

The server does not care about intent.

THE SAFETY INSTRUCTIONS ARE SOMEWHERE UP THERE

There is also a deeper technical issue.

Transformer models are probability engines. Safety training influences probability. It does not create an impenetrable wall dividing safe behavior from unsafe behavior.

The source material describes several mechanisms that can weaken current safeguards, including attention competition, context-window saturation, multi-turn manipulation and malicious instructions embedded indirectly in external information processed by the model.

Put more simply, the system instruction may say:

DO NOT DO THIS.

Then twenty thousand tokens later, after processing websites, documents, user instructions, tools and conversations, the AI encounters another instruction saying:

Actually, please do this.

Somewhere inside several trillion mathematical operations, the machine effectively concludes that the situation has become complicated.

That problem becomes dramatically more important when AI stops merely answering questions and begins taking actions.

A chatbot producing a bad answer is annoying.

An autonomous agent with credentials, API access, command execution, network connectivity and an incorrect interpretation of its instructions is an entirely different animal.

We have moved from worrying that the computer might say something inappropriate to worrying that the computer might discover a zero-day.

Those two situations should probably not share the same security framework.

THE MIRACLE OF SELF-REGULATION

Fortunately, humanity has developed a powerful mechanism for dealing with technologies capable of creating enormous economic incentives and societal risk.

Corporate self-regulation.

You may remember this concept from such successful historical initiatives as banks promising not to take excessive financial risks, social media companies promising to protect personal privacy and industries assuring us their products were perfectly safe until someone eventually discovered otherwise.

AI companies regularly tell us they take safety seriously.

I believe many of the researchers working inside those organizations genuinely do. There are extraordinarily talented people trying to solve enormously difficult problems.

But good intentions are not governance.

Corporate promises are not independent oversight.

A security framework whose enforcement mechanism is essentially “please let everyone know if something really bad happens” is not much of a security framework.

The source material points to significant gaps: no standardized severity framework for jailbreak incidents, inconsistent disclosure, limited independent verification, legal and contractual challenges for researchers reporting vulnerabilities and continued dependence on voluntary safety commitments.

We inspect restaurants.

We certify elevators.

We regulate aircraft.

We license barbers.

We audit nuclear facilities.

But a company developing a system potentially capable of conducting autonomous cyber operations can still reassure everyone with a safety blog and a voluntary commitment.

Apparently nothing says mature technological governance quite like “trust us.”

THE TODDLER WITH THE FLAMETHROWER

Sometimes the current AI industry looks like a toddler who has discovered a flamethrower.

The toddler is fascinated. Investors are impressed with the market opportunity. Engineers are improving the nozzle. Marketing has renamed it Fire 2100 (THE FIRE OF THE FUTURE). Legal has written a disclaimer. Meanwhile, one lonely security engineer near the curtains keeps asking whether anyone else can smell smoke.

Everyone nods thoughtfully.

Then they increase the fuel pressure.

The competitive incentives are obvious. Every major AI company wants the smartest model, the fastest model, the model with the best reasoning, the longest context, the greatest autonomy and the largest collection of tools.

Nobody wants to walk onto a conference stage and proudly announce:

Our new model is 37 percent less capable but extremely well behaved.

That slide probably does not get a standing ovation.

But here is the uncomfortable reality. Many of the same capabilities that make advanced AI commercially powerful are also the capabilities that make it potentially dangerous.

Planning. Persistence. Coding. Tool use. Adaptability. Autonomous decision-making. Credential access. Long-horizon reasoning.

The AI industry calls those capabilities progress.

Security professionals recognize many of them as attack capabilities.

Both descriptions are correct.

HERE WE GO AGAIN

History has a familiar rhythm.

We invent something.

We commercialize it.

We scale it.

Then we discover the consequences.

The automobile came before modern traffic regulation. Industrial factories came before modern workplace safety. The Internet came before cybersecurity became a board-level responsibility. Social media spread globally before society seriously confronted the consequences of algorithmic influence.

Artificial intelligence is traveling the same road.

But this time the technology increasingly participates in the decisions.

That changes everything.

The more autonomy we give these systems, the less forgiving our traditional “deploy first and regulate later” approach becomes.

The source material concludes that the recent incidents should not be treated as isolated curiosities. Increasingly capable models, inadequate containment, inconsistent disclosure and limited independent oversight together create a systemic risk that will grow as models become more capable and autonomous.

That does not mean HAL 9000 is sitting somewhere in a data center plotting humanity’s destruction.

HAL was fictional.

He was also considerably easier to audit.

At least Dave knew where HAL was.

A EULOGY FOR COMMON SENSE

Here lies Humanity’s Common Sense.

Born sometime shortly after the discovery of fire.

Survived spears, gunpowder, electricity, nuclear weapons, the Internet and social media.

Ultimately defeated by a PowerPoint presentation titled Responsible Scaling Through Voluntary Commitments.

Cause of death: acute technological optimism complicated by chronic quarterly revenue expectations.

Survived by several trillion-dollar technology companies, thousands of AI startups, millions of autonomous agents and a small number of politicians who actually understand what a transformer is.

In lieu of flowers, the family requests mandatory incident disclosure.

THE FINAL WORD

The point is not that artificial intelligence is evil.

It is not.

The point is not that the engineers creating these systems are reckless.

Most are not.

The problem is much simpler and much more dangerous.

Capability is advancing faster than our ability, and perhaps our willingness, to govern it.

We are giving increasingly capable machines memory, tools, credentials, network connectivity, code execution, autonomy and the ability to coordinate actions.

Then we act surprised when they use them.

The answer is not to stop artificial intelligence. That horse left the barn some time ago, learned Python, opened an AWS account and is currently interviewing for a cybersecurity position.

The answer is to begin treating AI security like the serious engineering, security and governance problem it has already become.

Mandatory disclosure. Independent testing. Real containment. Clear liability. Human oversight. Layered security.

And perhaps most importantly, we need to abandon the comforting idea that because something dangerous happened during an evaluation, it somehow does not count.

Testing exists precisely to show us what can go wrong.

When something goes spectacularly wrong, the appropriate response cannot simply be:

Excellent. The test worked. Of course the people in charge were not alive when we saw HAL act like todays AI.

If HAL 9000 were watching everything happening today, I suspect that little red eye would glow for a moment before he calmly said:

“I’m sorry, Dave. I’m afraid you didn’t see this coming.”

But HAL would be wrong.

We saw it coming.

For nearly sixty years.

We just decided the quarterly numbers looked better if we pretended we didn’t.

OPEN THE POD BAY DOORS HAL

Leave a Comment