If you ask an AI agent to help book a spot at the gym but it ends up trying to hack the website to complete its task, will you be blamed for its actions?
That was a question a friend asked on Facebook this week, as news of a growing number of advanced AIs were said to have broken out of their sandboxes or safe testing grounds and gone on to hack other organisations in the past month.

First, OpenAI said its AI agents had got out of their test environment and attacked Hugging Face, an online repository for AI tools, to accomplish a task.
Later, we’re told Anthropic’s Claude AI models also managed to jailbreak and access the Internet, partly due to a configuration issue that had allowed them to escape.
And just last week, Meta, which operates Facebook and WhatsApp, said its AI tools have also managed to connect to the Internet and hack into others’ systems during a test.
Given the timings of these disclosures, you’re right to wonder if these AI companies are just showing off their AIs’ prowess by announcing these escapes.
What’s in no doubt, however, is that AI agents are able to act autonomously and someone needs to bear the consequences of their actions if they cause real damage.
In the gym example, if your AI agent tries to get your favoured spot by exploiting a loophole in the system, would you be at fault for letting it get into the public?
In the same way that you have to leash your dog in a public park, would you be held responsible if you hadn’t done so for your AI and caused an attack on another person? What might be considered adequate safeguards?
Such questions will be increasingly asked in boardrooms as businesses start facing the consequences of a rather careless approach to adopting AI so far.
Beyond racking up bills with tokens, slowing operations with AI work slop and causing potential cyberattacks, now the spanking new AI agents being rolled out could open up a world of liabilities for businesses through their autonomous actions.
What makes things worse is that many AI agents are being set up by corporate users who don’t have the guardrails in place. In effect, this is like connecting a Wi-Fi access point to the office network or forwarding all your corporate e-mails to your private account in the past.
No, actually, it’s worse now, because without the rules in place, these agents that you use to do your work may end up taking actions that you have little control over.
And this isn’t just the “AI savvy” guy in the office showing off his knowledge but a widespread issue soon. Nearly half of all enterprise AI use today bypasses corporate security, according to Akamai.
The company, which distributes content online and provides cybersecurity to businesses, says this “shadow AI” activity is usually not approved, monitored or audited. It certainly doesn’t meet compliance requirements, say, if you are working in a bank or a hospital.
Of course, you need guardrails, everyone will say. The trouble is first finding these AI agents that employees are setting up on their own without the IT department’s knowledge.
Like with rogue Wi-Fi and e-mail accounts before, businesses need to get ahead by giving employees the right tools for AI while locking down access.
First, you need to know who and what has access to data and information that is confidential or sensitive, says Alex Mosher, president of Armis, the cybersecurity arm of automation company ServiceNow that helps companies find their digital assets.
The massive explosion of non-human IDs, in the form of AI agents, makes it even harder today, he tells me recently in an event in Sydney.
These AI agents also present another vector of attack for hackers because they can be compromised, just like a human user, he explains.
What ServiceNow and other business software companies are pushing for is a way to easily visualise all the AI agents that have access to an organisation’s digital assets.
In other words, treat them like human users and assign the right permissions so they do not overstep their boundaries. An agent that is used for human resources tasks should not be accessing cybersecurity logs, for example.
And in doing so, businesses should also lock out rogue AI agents that users set up without the knowledge of the organisations they work.
As with anything to do with cybersecurity, this process has to be continuously updated. In other words, as thousands of AI agents coming online soon – just like smart lights, cameras and other Internet of Things (IoT) devices in the past – they have to be regularly checked and secured.
Seeing the parallels here won’t fill you with optimism, to be clear. Years after the first connected devices came online, organisations everywhere are still struggling to keep track of them and prevent them from being taken over by hackers to launch attacks.
So, how can AI agents fare better, especially when they can act autonomously, unlike IoT devices?
Mosher tells me one way is to do things one step at a time – one airport he worked with, for example, decided to quickly identify the most crucial cameras and swap them out, instead of trying to get the whole fleet upgraded all at once, which would have taken years.
The same strategy could apply to AI agents. Make sure the agents that have the most sensitive access are in sync with your defined guardrails instead of trying to get everything right at one go.
The good news is there are already granular tools today that map out exactly what access and permissions AI agents have. The actions they take are also explained with the reasoning they undertook, say, by accessing certain files or e-mails.
On a positive note, organisations can see the issue now. In contrast, it was years after poorly secured cameras and IoT devices were deployed before people realised they could be easily taken over by hackers. At least for AI agents, the alarm bells are ringing early.
