Google's SynthID watermarks are secret trackers, argues new 'spymark' essay
Brandon Thomas published an essay titled "Spymarks, Not Watermarks," arguing that a category of technology already being deployed by major AI companies, exemplified by Google's SynthID, goes well beyond traditional watermarking and needs its own name: spymark. He defines a watermark as a visible mark embedded to verify authenticity or assert ownership, and a spymark as a hidden signal that makes a person's work traceable without their knowledge or consent.
According to Thomas, Google SynthID embeds secret signals, described by Google itself as "imperceptible to humans," into images, audio, text and video. He writes that this signal can encode database identifiers tied to a person's identity, including user records, full name, IP addresses, date of birth, physical addresses and political party affiliation. He cites Google's SynthID-Image paper, which reports that the SynthID-O variant can encode a 136-bit payload into a 512x512 pixel image, enough room for a 64-bit database identifier with 72 bits left over for error correction. He states that SynthID is not the only such system: OpenAI and other tech companies are developing similar tracking mechanisms at scale. He argues that while these companies frame the technology as a way to identify AI-generated content, they have built in tracking capability that goes beyond that stated purpose.
Thomas extends the argument to audio and text. He describes audio spymarks that make small changes to a waveform's time or frequency domain to encode identifying data, illustrating the point with spectrograms from a 2025 paper, "SoK: How Robust is Audio Watermarking in Generative AI models?" by Wen et al. He points to audiowmark, an open source tool that has existed since 2018 and can hide a 128-bit payload in audio while protecting it with a secret AES key. For text, he writes that SynthID steers word choices to create a statistical pattern that can itself encode a tracking payload.
The essay draws a distinction between spymarks and standard metadata such as EXIF tags in photos or ID3 tags in MP3 files, which he says are documented fields that a user can inspect and strip out. Spymark signals, by contrast, are invisible and can persist even after files are edited or their known metadata is removed. He cites printer tracking dots, in use since the 1980s, as an early real-world example of a spymark, and quotes Ursula K. Le Guin's line "To speak the name is to control the thing" as the reasoning behind coining a sharper term.
Thomas warns the risk is not theoretical: he raises the prospect of a future in which every device is attested and every social media post carries an account-linked spymark, making it possible to reconstruct how a piece of content, or the people connected to it, spreads online. He argues this would be especially dangerous for whistleblowers or anyone at risk of being persecuted for what they publish, and closes by calling a spymark "a spy tool used to spy on you and everyone you interact with."
Key facts
- Google SynthID embeds hidden signals, described by Google as 'imperceptible to humans,' into images, audio, text and video, which the essay says can encode database identifiers linked to a person's identity.
- Google's SynthID-Image paper says the SynthID-O variant can fit a 136-bit payload into a 512x512 pixel image, room for a 64-bit database identifier plus 72 bits of error correction.
- The open source tool audiowmark, in use since 2018, can hide a 128-bit payload in audio and lock it with a secret AES key.
- Unlike standard metadata such as EXIF or ID3 tags, which users can inspect and strip, the essay says spymark signals stay invisible and can survive metadata removal and some edits.
- Author Brandon Thomas proposes 'spymark' as a term to separate benign, visible watermarks from hidden tracking signals, citing printer tracking dots from the 1980s as an early precedent.
Why it matters
The essay reframes a technology already rolling out across AI-generated media, Google SynthID and, per the author, comparable systems from OpenAI and others, as a tracking mechanism rather than a content-authenticity feature. Its argument is terminological: renaming the hidden, tracking-capable variant of "watermark" to "spymark" is meant to make the privacy stakes legible to non-technical readers, similar to how "spyware" once reframed software that phoned home. That matters because these signals are already embedded in AI-generated images, audio, video and text with little public scrutiny of what they can encode.
Who it affects
Anyone whose photos, audio, video or AI-generated text passes through a system using SynthID or a similar tool, according to the essay. The author singles out whistleblowers and anyone who might face reprisals for what they publish as facing the sharpest risk, since a spymark could in principle let a tracked copy be traced back to the account or device that produced it.
How to use it
The essay's practical upshot is vocabulary rather than a tool: it asks readers to stop lumping tracking signals in with ordinary watermarks and to reserve "spymark" for the hidden, tracking kind. It also draws a concrete line for anyone trying to protect their files: standard metadata such as EXIF or ID3 tags can be inspected and stripped with ordinary tools, but the essay says spymark signals embedded in pixels, audio or word choice do not go away with a normal metadata wipe.
How solid is it
The technical claims are grounded in a cited Google research paper (SynthID-Image) for the 136-bit payload figure and in a 2025 academic paper, "SoK: How Robust is Audio Watermarking in Generative AI models?" by Wen et al., for the audio-watermarking material. The existence and 2018 origin of the open source audiowmark tool is independently checkable. The broader claim that OpenAI and unnamed "many other tech companies" are building similar systems "at scale" is the author's own assertion and is not sourced in the essay to a specific company statement or document.
Risks and caveats
This is an opinion and terminology piece by an independent author, not a peer-reviewed study or a report of a documented case where someone was actually identified through a watermark. It describes technical capacity, how many bits a signal could carry, rather than confirmed real-world identification of a specific person. The essay does not name which companies beyond Google and OpenAI are involved, does not date SynthID or its paper, and does not describe any regulatory or company response to the "spymark" framing.
“A spy tool used to spy on you and everyone you interact with.”
— Brandon Thomas, author of the essay