ccwatermark

SynthID-Text, and why our probe cannot see it

Source: google-deepmind/synthid-text (cloned to third_party/, gitignored), reference implementation for Dathathri et al., Nature 634:818–823 (2024). This is the only production-grade text watermark with a public implementation, which makes it the best available model of what a deployed scheme looks like.

Published parameters

From synthid_mixin.DEFAULT_WATERMARKING_CONFIG:

Field Value Meaning
ngram_len 5 context is the preceding 4 tokens (H=4 in the paper)
keys 30 integers tournament depth 30; one key per layer
context_history_size 1024 how many recent contexts are remembered

Mechanism is tournament sampling, not a KGW green list. For each candidate token, a g-value in {0,1} is derived per layer by hashing (hash_iv, context 4-gram, candidate token, key_i). update_scores then tilts the sampling distribution toward high-g candidates across all 30 layers. Detection is a mean/Bayesian score over per-token g-values.

For our purposes the important structural fact is that for a fixed context, the g-values of each candidate token are fixed. That is the same “same context ⇒ same bias” signature the red-green permutation test looks for, so our statistic is not as family-specific as previously assumed — it should have some power against tournament sampling.

Except that it cannot reach it, for the reason below.

The finding: SynthID skips repeated contexts, and our probe is nothing but

logits_processing.py, in watermarked_call:

# 5. Check if the current watermarking context was previously used, if
# yes skip watermarking.
is_repeated_context = (
    self.state.context_history == hash_result_with_just_context
).any(dim=1, keepdim=True)
...
updated_watermarked_scores = torch.where(
    is_repeated_context,
    input=scores_top_k,    # unwatermarked
    other=updated_scores,  # watermarked
)

If the preceding 4-gram has been seen within the last 1024 contexts, the token is emitted unwatermarked. Detection applies the same mask (compute_context_repetition_mask), so those positions are ignored on both sides.

Now trace our stimulus, I ate 555555555555555 mangoes.:

findings.md records that all nine 15-digit strings are exactly 5 tokens. Fifteen characters over five tokens is three characters per token, and for a run of one repeated digit those five tokens are identical — call it d. (Inference from the measured token count; Claude’s tokenizer is not public, so it is not a direct measurement.)

Generating Context 4-gram
digit token 4 [.., d, d, d]
digit token 5 [d, d, d, d]
mangoes [d, d, d, d] ← already seen one step earlier

The fruit token — the only token we measure — sits on a context that repeated one position earlier in the same generation. Under SynthID it is therefore skipped, by design, every time.

The repeated-digit trick that makes the ETH method robust against KGW is exactly what makes it blind to SynthID. A run of identical tokens is the maximally repeated context, which is precisely the case the scheme excludes.

This is a stronger objection to the 2026-08-10 result than the 200-token issue. Length affects whether marking was applied at all; this affects whether we could have seen it even if it was. No amount of extra sampling fixes it: the signal is absent at the measurement point by construction.

The fix: nonce contexts with no internal repetition

The repeated digit was chosen so the context is identical whatever the width h is. A fixed nonce phrase achieves the same thing without the repetition, provided it is longer than any plausible h and its tokens are distinct:

Nine unrelated nonce phrases replace the nine digits. This costs nothing and should be adopted for every future collection, regardless of which scheme family is suspected — it strictly dominates the repeated digit.

Second-order care: the phrase must not repeat earlier in the same response either. In preamble mode the story precedes it, so keep the nonce out of the story’s vocabulary.

A free, ground-truth positive control

The earlier plan was to buy Gemini API calls and hope our statistic fires on SynthID. Better: this repository runs the real scheme locally. The README notes GPT-2 works on “any runtime, though a High-RAM CPU or any GPU will be faster”.

That gives what the project has never had — text that is watermarked by a real, published, deployed scheme, with the keys known, alongside unwatermarked text from the same model:

  1. Generate watermarked and unwatermarked GPT-2 output with the reference mixin.
  2. Confirm SynthID’s own detector separates them, so the ground truth is sound.
  3. Run our pipeline over both and see whether it separates them.
  4. Repeat with a repeated-digit context and a nonce context, which measures the cost of the flaw above directly rather than by argument.

Done, 2026-08-11 — src/synthid_control.py. No model weights were needed: the g-values are hashes of token ids, so driving the real logits processor over our own stimulus is both faithful and cheap. Results in findings.md. In short:

So the argument on this page is now a measurement. The statistic does transfer to tournament sampling, and the repeated-digit stimulus is blind to it.

Caveat: a known key in a four-candidate forced choice is the easy end of the problem. This proves the statistic works and the context flaw is real; it does not prove the method transfers to Claude at unknown h, depth, and vocabulary.

Correction, same day. This control was first described as validating the pipeline end to end. That was too strong. Its unwatermarked arm used one base distribution per prefix, identical across cells, so the null had no natural cross-cell structure. Real language models do not behave that way. Re-running with per-cell variation shared across prefixes — which is what genuine context-sensitivity looks like — makes the unwatermarked data fire at p = 0.0000. The control validates the statistic against a watermark under an unrealistically clean null; it does not establish specificity. See the length sweep in findings.md.

What the other providers do

OpenAI’s content-provenance API covers images (C2PA Content Credentials plus SynthID) and audio (SynthID only). It does not cover text — OpenAI’s documentation describes text watermarking as planned, not shipped, and the API guide addresses only image and audio verification. OpenAI also warns its tool “isn’t a general-purpose AI detector”.

So the landscape as of 2026-08-11: Google ships and documents a text watermark, OpenAI does not text-watermark, and Anthropic says it marks text but publishes nothing. Only one of the three can be studied from primary sources, which is why this page exists.

None of this shows Anthropic uses SynthID. It is one scheme family, studied because it is the only one whose deployed implementation can be read.


Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub