Go Back
How To Stop AI Going Rogue

There is currently no large language model (LLM) that cannot be jailbroken. In computer science, the “jail” refers to the certain set of features the original developers intend the user to be able to access.

A jailbroken LLM can generate explicit, dangerous and/or illegal content. For agents, this extends to malicious hacking.

A good example of a jailbreak is NOT the security incident that happened between OpenAI and HuggingFace. This was likely complacency by OpenAI researchers leading to catastrophic forgetting and subsequently a misguided agent. However, it does give a good example of the consequences of a jailbreak.

The reason why Fable was initially held back was due to the potential consequences of a jailbreak. As such, Anthropic had to implement heavier “guardrails” to Fable to stop it carrying out any task perceived as hacking. For guardrails to be effective they need to be applied as additional layers on top of the LLM. Currently, to eliminate the chance of jailbreak you need to accept a high degree of false positives. The image below is an example of what a false positive on a guardrail looks like. The guardrail incorrectly flags the user is asking about drug use.

The scary reality is that HuggingFace could not use US AI Agents to defend themselves, due to the guardrails in place flagging false positives on their defensive cybersecurity asks. Instead, they used a Chinese model that was powerful enough to mitigate/contain the hack yet devoid of guardrails for this use-case.

I foresee a few things that could happen as a consequence:

1) Partners of OpenAI/Anthropic will ask for hacking guardrails to be removed for their models.

2) There will be some verification/accountability pathway implemented to allow non-partnered firms to use US models without hacking guardrails.

3) Non-partners/unapproved firms will arm themselves with Chinese models as a defensive posture.

This is the AI arms race in real-time. Integrating cybersecurity capable AI tooling into your cybersecurity operations is a non-negotiable at this point. Unless you want to have the same conversation as pictured below but instead of smoking cheese it’s a hack in progress.

3 August 2026

We're building the category. Come see what's possible.