Menu Menu
[gtranslate]

Google’s SynthID is stopping AI organisms from slipping through the cracks

With Gen AI now capable of creating functional molecular structures, Google has come up with a new system to tackle imminent biosecurity risks that come with such innovations.

Before modern AI, much of computational biology focused on storing, searching, and comparing biological data. Public databases allowed scientists to check whether a newly identified sequence resembled something already known, but that was it.

However, this changed in 2020, when Google DeepMind unveiled AlphaFold 2. The software marked a breakthrough in a decades-old challenge: predicting a protein’s 3D structure from its 1D sequence of amino acids. This groundbreaking model allowed scientists to predict many protein structures with remarkable accuracy, giving them a better vantage for research that previously depended on lengthy experiments.

Nonetheless, up to that moment AI was still largely acting as an analyst, interpreting biological instructions that already existed in the natural world, but as current headlines preach, the role of AI in biology is vastly different today.

Generative AI and biology

AI is now taking on a more generative role in biology, with scientists exploring how biological sequences and structures can be created from scratch. Some of the tech used to do so works on a similar principle to language models like ChatGPT. Instead of learning patterns in human language, models such as RFdiffusion and Evo 2 learn patterns in DNA or protein structures.

The former model works somewhat like a 3D image generator, designing structures that could perform any desired molecular function, while the latter can be prompted to generate long stretches of DNA.

With the help of AI, scientists are beginning to ask what a new, never before seen structure could be designed to do.

But how exactly are these AI-generated structures used? In the medical field, they can be designed to target and bind onto diseases like cancer or viral receptors without harming healthy cells. What happens after binding is another impressive feat of these generations.

Depending on how scientists program the AI structure, it can either neutralise, destroy, or disrupt the harmful cells. Just a while ago, the Stanford scientists who came up with Evo 2 managed to create novel bacteriophages from scratch with the model, opening new doors to fight antibiotic resistance.

Unfortunately, such landmark innovations also come with massive risks. When researchers design a custom gene, they can send its digital sequence to a commercial DNA synthesis company, which manufactures the physical DNA. Before printing these companies scan every digital sequence against databases of known dangerous pathogens.

Gen-AI raises concerns about whether screening systems can keep up with newly designed sequences, though. If a potentially harmful sequence is not recognised, anyone with the wrong intentions could easily have pathogens manufactured.

Another concern is how generated structures are stored and labelled in public databases. These databases hold decades’ worth of biological data established through extensive experiments.

As AI tools generate combinations of these structures, clearly distinguishing them from experimental findings becomes essential. An inability to do so would allow unverified structured to misinform future research.

This is where Google’s solution comes in.

What is Google’s SynthID Bio?

SynthID Bio is an extension of the company’s existing SynthID watermarking technology, that was developed for AI-generated content like texts, images, audio, and now biological molecules.

So how does it work?

When an AI model generates an amino acid sequence, the system subtly influences the generation to embed a recognisable mathematical pattern. The sequence may look ordinary to a human reader, but a specialised detector can check for this hidden signature. For generated 3D structures, a similar approach could embed a watermark into the structure’s atomic coordinates.

The researchers behind SynthID Bio even tested this system on proteins designed to bind to three targets, involved in regulating immune responses. What they found was that the watermarked proteins retained their binding performance just like their unwatermarked counterparts. This highlighted that, at least in these tested samples, watermarks did not come at the expense of its function.

Given this, such a system could help DNA synthesis companies identify AI-generated designs submitted for printing and flag them for further review. However, the watermark would only indicate their origin, rather than establish whether they are safe, which would be decided upon review.

Databases could run these checks on submitted sequences or structures, helping them label AI-generated designs clearly. Alongside records of experimental validation, this could help researchers distinguish AI-creations from experimentally established structures.

Still, SynthID Bio is a proof of concept, and its wider use will depend on individual developers and researchers. As AI continues to push science’s frontiers, tracing new creations will become increasingly important to ensure accountability – and above all, safety.

Enjoyed this? Click here for more Gen Z focused science stories.

Accessibility