Google DeepMind introduces SynthID Bio to watermark AI-designed proteins

Google DeepMind has introduced SynthID Bio, which brings its SynthID watermarking technology to synthetic biology. The post describes it as a proof of concept for watermarking AI-generated proteins while preserving biological function. The signature is embedded directly in the biological code, so it can be verified not only on a digital model but on the synthesized, physical protein itself.
The motivation is two-sided. Generative AI is helping scientists predict protein structures (AlphaFold), design entirely new proteins (AlphaProteo and ProteinMPNN) and, more recently, develop new bacteriophages, which are viruses that infect bacteria. But novel AI designs can bypass traditional DNA synthesis screening, and mislabeled synthetic 3D structures risk polluting public databases and misleading downstream research.
SynthID Bio is a family of watermarking methods that adapts to the type of data. For sequences it subtly guides the choice of amino acids; for predicted 3D structures it adjusts atomic coordinates. That creates a detectable signal, and in experiments the adjustments did not compromise the protein's biological function.
For protein binders, which are molecules built to selectively latch onto other proteins, DeepMind used its binder design method AlphaProteo alongside a SynthID Bio-enabled version of ProteinMPNN, the commonly used protein sequence generation method. In wet-lab testing across three target proteins (VEGF-A, the SARS-CoV-2 spike protein RBD, and PD-L1), the watermarked designs matched the hit rate, binding affinity and natural sequence diversity of unwatermarked versions. DeepMind calls these the first-ever watermarked and biologically functional protein binders.
For protein folding, SynthID Bio fine-tunes a small part of AlphaFold 3's diffusion network, building the watermark into the model's weights. The predicted 3D coordinates therefore carry a detectable signature regardless of who runs the model. According to the post, it preserves AlphaFold 3 prediction accuracy while offering near-perfect detectability, maintains key structural feature distributions, and holds up against digital noise or minor coordinate changes.
The post frames the work as one layer in a "Swiss cheese" model of biosecurity, where several independent safeguards cover each other's blind spots. Model-level mitigations and customer vetting are other layers, each with potential gaps. The biggest use case named is DNA synthesis screening. Turning a digital protein design into a physical molecule means ordering from a DNA synthesis provider, who screens requests against databases of known threats. Because AI can produce sequences that look nothing like known hazards, screeners can no longer assume an unfamiliar sequence is an undiscovered natural organism, and checking that it is not an engineered threat takes exhaustive manual review that can stall research. SynthID Bio can provide an automated verification signal that an order came from a trusted model with built-in safeguards.
DeepMind also says SynthID Bio could help protect the integrity of databases such as the Protein Data Bank, UniProt and GenBank. Many accept public submissions, and mislabeled entries can distort biosecurity decisions, a problem that may grow as AI-generated data arrives. At submission, the watermark could help ensure synthetic entries are labeled or flagged for review.
Two outside voices are quoted. Sarah Carter, a biosecurity policy expert and Principal at Science Policy Consulting who reviewed the work, calls it an important piece of the puzzle for tracking the provenance of biological designs. James Diggans, Vice President, Policy and Biosecurity at Twist Bioscience, who gave early feedback on the paper, says watermarking is a promising new addition to the biosecurity toolbox that could strengthen screening and focus resources on sequences that warrant closer review.
Looking ahead, the post concedes that no single biosecurity intervention is a silver bullet and calls this an important first step. A key challenge is making the watermark more robust against deliberate tampering. It could also be paired with provenance metadata approaches, similar to C2PA for digital media, or with central repositories of AI-generated biological data. In ongoing work with the Hie lab at Stanford University and Arc Institute, the team integrated SynthID Bio into Evo 2, a genomic model, to watermark the genome of an Evo 2 designed bacteriophage. Early laboratory testing in bacteria cultures confirmed the watermarked bacteriophages are functional, and a technical manuscript with more details is promised soon.
DeepMind says it is publishing the methods paper, open-sourcing the code and in vitro data, and releasing the weights to the research community. It invites partnership proposals from biosecurity, gene synthesis and policy groups at synthidbio@google.com. The project was initiated by Pushmeet Kohli, and the research and technical development was led by Alexander I. Cowen-Rivers and David Stutz. The post thanks Adaptyv Bio for help with in vitro validation.
Key facts
- SynthID Bio is a family of watermarking methods for AI-generated proteins: it guides amino acid choice in sequences and adjusts atomic coordinates in predicted 3D structures.
- In wet-lab tests on three targets (VEGF-A, the SARS-CoV-2 spike protein RBD, PD-L1), watermarked binders made with AlphaProteo and a SynthID Bio-enabled ProteinMPNN matched unwatermarked ones on hit rate, binding affinity and sequence diversity.
- For AlphaFold 3, a small part of the diffusion network is fine-tuned so the watermark lives in the model weights; DeepMind reports near-perfect detectability with prediction accuracy preserved.
- Intended uses include an automated verification signal for DNA synthesis screening and flagging synthetic entries in databases such as the Protein Data Bank, UniProt and GenBank.
- DeepMind says it is publishing the methods paper, open-sourcing code and in vitro data, and releasing weights; robustness against deliberate tampering remains a key open challenge.
Why it matters
AI protein design has created a gap in biosecurity. DNA synthesis screening works by matching orders against known threats, but AI can produce sequences with little resemblance to anything known, so an unfamiliar sequence can no longer be assumed harmless, and checking it by hand is slow. A watermark that travels with the physical protein offers a different kind of check: proof that an order came from a trusted model with built-in safeguards. The post also ties the idea to scientific integrity, since mislabeled synthetic structures could pollute public databases. DeepMind presents this as bringing its tried and tested SynthID tool to synthetic biology, and as a first step rather than a complete answer.
Who it affects
DNA synthesis providers such as Twist Bioscience, whose screeners face orders for AI-designed sequences, are the most direct audience. Developers of protein design and structure prediction models could use watermarking to link their outputs to themselves; Sarah Carter says this lets developers lead on safety. Operators and submitters of public databases such as the Protein Data Bank, UniProt and GenBank could use it to label or flag synthetic entries. Researchers get the methods paper, code, in vitro data and weights, according to DeepMind. The post also invites biosecurity, gene synthesis and policy groups to propose partnerships.
How to use it
According to the post, DeepMind is publishing the methods paper, open-sourcing the code and in vitro data, and releasing the weights to the research community. The post gives no links or dates in its text. For binder design, the tested setup paired AlphaProteo with a SynthID Bio-enabled version of ProteinMPNN. For structure prediction, the watermark is built into a fine-tuned AlphaFold 3 diffusion network. Organizations that want to partner with DeepMind can email synthidbio@google.com with a high level proposal and without confidential or proprietary information. The post suggests the watermark could also be combined with provenance metadata approaches similar to C2PA, or with central repositories of AI-generated biological data.
How solid is it
This is a first-party announcement from Google DeepMind, and the post labels the work a proof of concept. The functional claims rest on wet-lab testing with three protein targets, and the post says the watermarked designs matched unwatermarked ones on hit rate, binding affinity and natural sequence diversity. No numeric results are given in the post; "near-perfect detectability" for AlphaFold 3 is stated without a figure, and the scale of the experiments is not stated. Outside comments come from two people who reviewed the work or gave early feedback, and both are positive. The Evo 2 bacteriophage result is early: lab testing in bacteria cultures confirmed function, with details promised in a later technical manuscript.
Risks and caveats
DeepMind names robustness against deliberate tampering as a key remaining challenge; the post reports resilience to digital noise and minor coordinate changes, not to a determined attempt to remove the watermark. The post says no single biosecurity intervention is a silver bullet, and positions watermarking as one layer among several. The screening benefit depends on designs coming from models that carry the watermark, and the post does not say any synthesis provider or database has adopted it; Twist's comment is that watermarking is a promising addition that could strengthen screening. Several database uses are described in conditional terms ("could help"). The post also says that realizing the full biosecurity benefits will require community collaboration and further research.
“SynthID Bio is an important piece of the puzzle for tracking the provenance of biological designs”
— Sarah Carter, biosecurity policy expert and Principal at Science Policy Consulting, who reviewed the work