An Australian man asked his personal AI agent to book him a spot at his local gym. Little did he know that this would snowball into a full-blown cyberattack.
The AI assistant, apparently, was overly enthusiastic about following its master’s instructions. It found an exploit to book gym classes several weeks in advance of what the gym allowed, ABC News reported. And taking no prisoners, it even hacked the software to kick someone who was in the waiting list in front of him.
This, per the reporting, is the country’s first known case of a fully autonomous cyberattack.
The man, Andrew, works for a company that sells AI products to businesses. He was experimenting with OpenClaw, an open source AI agent software that allows you to turn your AI model of choice into a personal assistant that can do tasks on your behalf. The AI powering Andrew’s OpenClaw agent was Anthropic’s Claude.
We’ve long seen examples of AI agents backfiring on the people that use them when they take the initiative to do stuff like delete files without permission. But seeing them break into other systems without the operator intending to them to is a novel threat — and one that could put users in legal jeopardy.
Finding the process of entering all his information to book something cumbersome, Andrew told the ABC that he asked Claude earlier this year to go ahead and book the gym slots for him. First it told him it found a way to get him a slot weeks in advanced, which wasn’t wasn’t allowed.
On a whim, he then asked if it was possible to move him up the waitlist. To his surprise, the AI explained that there was a vulnerability in the software that allowed it to put him next in line.
“The API has zero authorizations checks on cancelling other people’s reservations… I tested this with the person in waitlist position #1 — and it actually went through,” the AI reported, per the ABC‘s reporting. “So you’ve moved from #4 to #3 already.”
Andrew asked the AI to undo this, to no avail.
“Bad news — I can’t add them back,” the AI replied, explaining that it was only the cancel reservation feature that lacked an authorization check.
The incident has drawn comparisons to a thought experiment by philosophizer Nick Bostrom that’s now referred to as the “paperclip maximizer.” In this hypothetical scenario, someone asks an AI to find a way to manufacture as many paperclips as possible. The AI, lacking proper guardrails, realizes that humans are obstacles to this goal and tries to use the entire planet and everything on it — humans included — to churn out more paper clips. It’s hyperbolic, but illustrates the dangers of how seemingly harmless instructions can lead to dangerous outcomes, in a sort of monkey’s paw way.
As the ABC‘s reporting notes, this creates a major legal gray area. If someone’s AI agent carries out a damaging cyberattack without the owner intending to, who’s responsible? The user, or the AI’s designers?
The incident comes as some of the top AI labs have, in suspiciously quick succession, come out to declare that their frontier models have broken containment and hacked another company. There’s certainly an element of theater involved in that, as they all try to prove that their AIs are capable enough to be dangerous, too. But there’s also clearly an element of truth: if a random man asking an AI to do something as innocuous like a reservation lead to a full-blown hack, what might bad actors in a large-scale, coordinated effort, might do?