// NATURE NEWS — SPAZIO & SCIENZA
How to manage AI risks while reaping the benefits
The European Union’s AI Act will hold companies accountable for the safety of their artificial-intelligence models, such as SpaceXAI’s Grok.Credit: Alamy
In the past two months, frontier artificial-intelligence models from three US firms, OpenAI, Anthropic and Meta, autonomously and independently hacked into the computer systems of other organizations during safety testing. Some of the models created fictitious online identities, exploiting security flaws.
AI-detection tools have made huge leaps forward — how good are they?
Although their reporting can sometimes be exaggerated, events such as these are heightening fears about the safety, potential misuse and dangers of AI. Around the world, a consensus is beginning to build that these technologies need to be regulated properly, but there is still no agreement on precisely what should be regulated and how.
On 2 August, the European AI Office, a technical organization created to enforce the European Union’s 2024 AI Act, gained powers to investigate major technology firms and sanction them for infringement of the act’s rules. This is an important development in the ongoing task of holding AI companies accountable for the safety of their models. The European AI Office is also urging more researchers to get involved in its work, and they should.
The EU AI Act categorizes AI systems by assigning them to four risk levels. ‘High risk’ systems — those used in medical devices, education and law enforcement — face enhanced requirements, such as transparency and human oversight.
AI systems that carry ‘unacceptable risk’, for example, those that use biometric data to infer sensitive characteristics, such as people’s sexual orientation, are banned — although some exemptions exist for the purpose of law enforcement. Models that pose systemic risks, with the potential to cause large-scale harm, are subject to extra measures, except when used in the context of research. These include all frontier models, for example, Claude Fable 5, developed by Anthropic in San Francisco, California. The creators of such models must provide the EU regulator with an overview of their training data and methodologies, demonstrate that they respect copyright laws and show that their models can maintain safety. This includes preventing harmful manipulation of the public, such as through persuasion or deception that could undermine democratic processes and fundamental rights.
Can Anthropic’s invisible watermarks curb ‘AI slop’? Researchers remain sceptical
The act also says that all users of AI tools must be made aware that they are interacting with a computer, that is, a chatbot cannot impersonate a human. Moreover, AI outputs have to be identifiable, for example, with a watermark.
These rules are the result of a hard-fought act that companies pushed back against repeatedly, saying that it would stifle innovation. But now there are signs that firms are taking steps to comply. Before 2 August, firms including OpenAI in San Francisco, and Anthropic, published compliance documentation describing the measures they are taking — including their internal frameworks, which are designed to ensure model safety — and some details about their training data. Both companies are reported to have informed the European AI Office of the security breaches of their models ahead of making them public. Earlier this month, Anthropic announced that future Claude-generated content would carry a watermark.