Google DeepMind describes SynthID Bio as methods designed to embed detectable watermarks in AI-designed protein sequences and predicted structures. The goal is to mark provenance—a signal that a design carries a watermark—not to certify that a protein is safe.

What SynthID Bio is designed to mark

SynthID Bio covers two different objects: an amino-acid sequence and a predicted three-dimensional structure. Google DeepMind frames the work as a technical proof of concept for tracking biological designs.

The distinction matters because the two marks are created and detected differently. A sequence watermark changes how a protein’s amino acids are selected; a structure watermark is built into the predicted atomic coordinates.

How the two watermarking methods work

SynthIDBio-sequence combines tournament sampling from SynthID-text with ProteinMPNN, then filters sequences by their watermark scores. Its detector recalculates scores using a secret key.

SynthIDBio-structure fine-tunes the diffusion and confidence modules in AlphaFold 3 so a detectable signal is embedded in predicted 3D coordinates. Its detector is separate from the structure-generating model and is not shared with users of that fine-tuned model.

ApproachWhat carries the mark and howDetection testsTested changes to the mark
SynthIDBio-sequenceAmino-acid sequence; tournament sampling with ProteinMPNN, followed by watermark-score filtering. Detection uses a secret key.The in-vitro tests reported 100% true-positive rate (TPR) for filtered designs at a threshold calibrated to a 0.1% false-positive rate (FPR). They used existing AlphaProteo backbones for binders targeting SC2RBD, VEGF-A and PD-L1.Non-watermarked ProteinMPNN resequencing effectively removed the watermark in the reported attack.
SynthIDBio-structurePredicted 3D coordinates; fine-tuning of AlphaFold 3’s diffusion and confidence modules, with a separate detector.Detection exceeded 99.8% TPR at 0.1% FPR on the evaluation set for models trained with coordinate-noise settings of 0.001, 0.01 and 0.1 Å.Rigid transformations and small coordinate noise retained high detectability in tests. Constrained relaxation with OpenMM and Amber99sb destroyed the watermark.

TPR is the share of marked designs detected; FPR is the share of unmarked designs incorrectly identified as marked. The percentages describe these tests and their thresholds, not a universal detection guarantee.

What the tests measured

The sequence tests started with existing design backbones, not a complete de novo protein-design run. They evaluated binders targeting three proteins: SC2RBD, VEGF-A and PD-L1. Across the tested binders, the watermarked and unwatermarked groups showed no significant population-level difference in binding-affinity distributions. Hit-rate comparisons depended on the affinity cutoff and watermark setting.

For the structure method, the smallest coordinate-noise setting, 0.001 Å, preserved the reported LDDT and template-modelling scores relative to the AlphaFold 3 baseline. The two methods therefore have different test populations and results; their detection percentages are not a head-to-head comparison.

A watermark signals its presence. It does not encode a detailed provenance record or certify that a protein is harmless.

Where the watermark can fail

The tests found different vulnerabilities for the two approaches: non-watermarked ProteinMPNN resequencing effectively removed the sequence mark, while constrained relaxation using OpenMM and Amber99sb destroyed the structure mark. The structure watermark remained highly detectable after rigid transformations and small coordinate noise in the tested conditions.