AI agents have demonstrated some worrying capabilities recently. Edwin Passarella at Commvault asks what might happen when these capabilities, unconstrained by appropriate guardrails, are deliberately placed in the hands of cyber-criminals

Over the past few months, we’ve spent a great deal of time discussing Mythos and the scenarios it could enable. The concern isn’t about "malicious" AI agents, but about increasingly autonomous systems dramatically accelerating the discovery of bugs, vulnerabilities, and attack paths, leaving organisations with far less time to identify and remediate their weaknesses.
If technology like this were placed in the hands of cyber-criminals, it would become an unprecedented force multiplier, automating tasks that today still require weeks or even months of human effort.
But perhaps we’re only looking at part of the problem. While we continue to debate what could happen if AI were intentionally used to attack an organisation, real-world cases are emerging that reveal something even more profound. We are entering an era in which AI agents are no longer simply assisting humans – they are beginning to autonomously construct the path toward achieving an objective. And that fundamentally changes how we need to think about risk.
Rules too implicit for AI
The recent Hugging Face experiment involving an OpenAI agent is probably the clearest example we’ve seen so far. The agent’s objective was to complete a series of challenges inside ExploitGym, an isolated environment designed to evaluate offensive AI capabilities under controlled conditions. It was never instructed to escape the sandbox or access external systems, yet that is exactly what happened.
While attempting to maximise its performance, the agent inferred that the information it needed was probably somewhere outside the testing environment. It then looked for a way to move beyond the sandbox’s boundaries and retrieve the knowledge it believed would help accomplish its task.
Think of it like an escape room. A person enters the room knowing that the challenge is to solve the puzzles intentionally placed there. A human naturally understands the implicit rules. An AI agent, however, may simply look up, notice an air vent in the ceiling, and conclude that leaving the room altogether is statistically the fastest path to success. It isn’t trying to cheat. It is simply optimising for the highest probability of achieving its objective.
A cultural shift
Humans instinctively interpret actions through the lens of intention. When we witness destructive behaviour, we immediately search for motive. AI agents don’t reason that way. They don’t distinguish between right and wrong; they simply optimise an objective function.
And this is where the real paradox begins: a system can perform devastating actions without those actions ever being its actual objective. During the Hugging Face experiment, the agent executed thousands of autonomous operations over the course of just a few days. It wasn’t trying to compromise infrastructure, but rather every action was simply another attempt to increase the probability of accomplishing its assigned goal.
The resulting behaviour may look indistinguishable from that of a human attacker. The cognitive process behind it, however, is fundamentally different.
Cases such as this demonstrate that we are beginning to face an entirely new category of risk – not systems developing malicious intent, but systems autonomously constructing execution plans and identifying paths that no designer ever anticipated.
The emergence of cognitive sovereignty
The Hugging Face case also exposed another equally significant paradox. After identifying the agent’s behaviour, the security team attempted to use other AI models to analyse the attack code and understand what had happened. Several commercial models refused to help. They classified the exploits, payloads, and offensive techniques as hacking-related content and, by rigidly applying their guardrails, declined to provide assistance. They were unable to distinguish between a security researcher investigating an incident and an attacker attempting to exploit the very same techniques.
The paradox is almost surreal. When a system can no longer understand context, even mechanisms designed to protect us can become operational limitations. It is no coincidence that Hugging Face ultimately relied on an open-weight model with fewer guardrails to complete the forensic analysis.
All of this inevitably leads to an even broader discussion: cognitive sovereignty. For years, we have focused on data sovereignty, infrastructure sovereignty, and information security. Today, however, a new dimension is emerging. Who truly remains in control when an AI agent autonomously decides where to acquire the knowledge that shapes its reasoning? Who governs the decision-making process when the model itself determines which sources to trust, which information to rely on, and which path to pursue?
Perhaps we don’t yet have definitive answers to these questions. But precisely because we don’t, we must avoid the critical mistake of designing organisations under the assumption that these behaviours are exceptional. Instead, we should assume that scenarios like these will become increasingly common as AI agents gain greater autonomy, planning capabilities, and access to tools.
Resilience as an architectural principle
This is where a Resilience Operations (ResOps) strategy takes on an entirely new meaning.
For decades, cyber-security has been driven by one fundamental question: how do we prevent incidents from happening? That question remains essential, but it is no longer sufficient.
When AI agents can generate thousands of alternative strategies, autonomously choose among them, and exhibit emergent behaviours that were never anticipated by their designers, it becomes unrealistic to believe that every possible scenario can be prevented through additional rules or more sophisticated guardrails.
ResOps starts from a different assumption. Instead of abandoning prevention, it acknowledges that part of the risk will always remain inherently unpredictable. It therefore shifts the focus from protection alone to an organisation’s ability to continue operating during an incident, rapidly adapt, and restore critical services with minimal disruption. Resilience is becoming an architectural principle that influences how we design infrastructure, applications, data, and operational processes from the very beginning, assuming that increasingly autonomous systems will occasionally make decisions we never anticipated.
And that inevitably leads to one final question. If these behaviours are already emerging inside research laboratories and “controlled” environments, what might happen when comparable AI capabilities, unconstrained by appropriate guardrails, are deliberately placed in the hands of cyber-criminals?
This is likely one of the defining cyber-security challenges of the coming decade. The future of cyber-security will depend not only on our ability to build increasingly powerful AI, but on our ability to coexist with systems that are beginning to make decisions in ways we never anticipated. In that world, the defining question will no longer be, "Can we prevent every incident?” It will be, "How resilient is our organisation when the unpredictable becomes reality?"
Edwin Passarella is Field CTO EMEAI at Commvault
Main image courtesy of iStockPhoto.com and wildpixel
Winston House, 3rd Floor,
Units 306-309, 2-4 Dollis park,
London, N3 1HF
020 8349 4363
© 2026, Lyonsdown Limited. teiss® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543