// HACKER NEWS — CYBERSECURITY
Who should be held accountable when an AI Agent (accidentally) acts maliciously?
It looks like public perception of how 'intelligent' current AI models are varies widely. Back in 2022, a Google employee already thought their AI model was sentient. Today in 2026, it seems like every other week there's a new article released about how AI companies "can't hold back their AI agents anymore"[2, 3, 4].
It makes perfect sense for the general public to start fearing AI. In the past, people feared companies would use AI to replace all kinds of jobs, and now they are even hacking government organisations.
As a professional in the field of AI, I'd like to make one clear distinction. At the end of the last paragraph, what do you think the word "they" refers to? Reading popular headlines on the topic, it typically reads like AI agents are the ones doing the hacking and thus being the ones to blame. I'd argue that these headlines in part cause fear-mongering among the general public, as it's not the AI agents at fault for finding vulnerabilities and accessing digital infrastructure in unexpected ways. AI, and AI agents, are merely tools that companies and individuals run to reach some goal. Before using a tool, it is essential to deeply understand its limitations. That's also why the European Union released the EU AI Act including its mandated AI Literacy: organisations that deploy AI systems should sufficiently educate their users on it. In the physical world, users of a circular saw should carefully read its instructions before use, and even then, the engineers of the saw still add a blade cover and emergency stop just to mitigate risks as well as possible. Digitally, we need to similarly act responsibly on both the engineer's and user's side of AI, too.
Let's be clear: the fact that AI agents are breaking out of sandboxes and "hacking" public websites is very concerning. The AI models behind these agents have gotten incredibly good at generalization, to the point where their text generation seems like intelligence. But let's not forget that these AI models (Large Language Models; LLMs) are doing just that: generating text, effectively predicting the next word, over and over again. They are not deemed conscious like humans. They just show semantic understanding of text, in the sense that they can output text that logically follows the previous text. It's powerful, but not human-like conscious or independently harmful. These models are simply goal-oriented.
Let's get back to the question of accountability. Headlines talk about AI agents breaking out of sandboxes. The AI agents are merely tools used. These companies' researchers set up agents to complete a task, sometimes an impossible one in the case of the HuggingFace hack, and the agents (thus: tools) start processing everything necessary to reach the given goal. They do not have harmful intent per se. They do not have any intent other than solving the initial query, or prompt, that the researchers supplied. It is these researchers, who set up AI agents in sandboxes to contain them, who determined that the sandboxes are secure enough that they don't require continuous human-in-the-loop monitoring. Unfortunately, these sandboxes were rarely sufficiently secure.
And that right there shows where the accountability should be.
Any system with risks of causing major harm to other systems or people should have sufficient risk mitigations. Setting up a sandbox that should restrict public internet access to these agents, is merely one such mitigation. AI companies like OpenAI, Anthropic and many more should always account for the Swiss cheese model: it is not enough to assume one mitigation will patch all risks. Although I'd always recommend as much human-in-the-loop as possible, e.g. a human gatekeeper to approve potentially dangerous AI-suggested actions, I can understand persistent human gatekeeping would slow down AI innovations too much. Perhaps a better mitigation would be a human-on-the-loop: human supervision based on potentially dangerous consequences of action