// ARS TECHNICA — INTELLIGENZA ARTIFICIALE
Google figures out how to watermark AI-designed proteins
Intended to help with biosecurity, it works with a popular AI protein design tool.
AI-based tools seem to be causing security threats on a nearly daily basis, in part because we’ve been slow to recognize potential threats. One area where we seem to be ahead of the game, however, is in biosecurity.
As with most things biological, utility goes hand in hand with threats. We’ve developed increasingly sophisticated tools for designing proteins and seen some major successes, such as AI-designed enzymes that can digest plastics or block venom proteins. But these same tools could be used to make toxins or alter the behavior of viral proteins.
And the software we use to identify DNA sequences that encode potentially threatening proteins doesn’t pick out AI-designed proteins, since nobody has characterized them well enough to know that they’re threats. Nearly a year after that risk was flagged, it still wasn’t clear what anyone could do about it.
On Wednesday, the DeepMind team at Google published a research paper offering a potential solution: protein watermarking. The system creates a watermark on protein sequences themselves without compromising the protein’s function. This allows new proteins designed by trusted researchers to be identified, opening everything else up to closer scrutiny.
The work was based on Google’s SynthID tech, which can add a subtle watermark to AI-generated digital material. The watermark influences the probability of certain choices the AI makes, and that bias ends up systematically distributed throughout the product, whether it’s text or images. Because you can’t identify the watermark without knowing how it was encoded, it’s impossible to remove. And because it’s distributed throughout the image, it can survive basic exporting, resizing, and so on.
It’s pretty easy to see how this can work with subtle differences in things like the colors of a photo. It’s a whole lot harder to see how you can do it with a protein.
Proteins are composed of only 20 amino acids, any of which could be essential for structural integrity or catalytic activity. While some of these amino acids are chemically similar (like leucine and isoleucine), others have opposite charges. Many proteins have significant regions where limited changes to their amino acid sequence are tolerable and other areas where even a slight deviation from the existing sequence inactivates the protein.
Proteins are also small. While images often contain millions of pixels, proteins containing 500 amino acids are fairly large. That’s a lot less raw material to hide any sort of signal in.
So it wasn’t clear the SynthID tech would work; it might be unable to hide sufficient signal in a typical protein, or, if it crammed in enough information to create a functional watermark, the resulting proteins might be inactive. The only way to find out was to try it.