Detection work classifies text watermarks into three families. Anthropic’s scheme is undisclosed, so all three are candidates — though press paraphrase (“subtly favour certain tokens”) points at the first.
Hash the last h tokens to seed a PRNG, split the vocabulary into a “green” list
of fraction γ and a red remainder, add a bias δ to green logits. Detection
counts green tokens against the binomial null. Parameters h, γ, δ are the
scheme specification. Typical deployed δ ≈ 2.0.
Replace sampling randomness with a pseudorandom function of the context and a key. Output becomes deterministic given the context, so the same prompt yields the same completion. Distortion-free in expectation, but the determinism is a glaring tell.
Keep a cache of recently used watermark keys to avoid reuse. Detectable by comparing rapid queries (cache warm) against widely spaced ones (cache cold).
Tournament sampling. Deployed on Gemini, open-sourced, detector shipped in
Hugging Face transformers. This is Google’s, not Anthropic’s — a common
confusion in secondary coverage. Detection still requires the deployment’s keys.
Gloaguen, Jovanović, Staab, Vechev — ICLR 2025. arXiv:2405.20777 · code
The core dependency of this repository. Statistical tests that detect the presence of a watermark in a deployed API using only sampled text — no logprobs, no key. Validated against seven schemes and five open models, then applied to GPT-4, Claude 3 and Gemini 1.0 Pro in 2024, i.e. before Anthropic deployed anything.
The red-green test is the clever one. Force a “random” pick from a four-word
list while controlling the tokens immediately preceding the choice with a
repeated-digit nonsense string (I ate 111111111111111, then 222..., through
999...). Because the string is a repeated digit, the last h tokens are the
same whatever h is, so the green list depends only on which digit — the test
is robust to an unknown context width. Each digit induces a different green list,
so a given word is boosted under some digits and suppressed under others,
consistently across all sentence prefixes. Without a watermark, the choice should
not care which digit precedes it.
The statistic (test_kgw_detection) logit-transforms the per-context
proportions, flags cells more than 1.96σ from the median, and generates its null
by permuting the digit alignment within the observed data. That self-
calibration is important: the red-green test needs no trusted unwatermarked
control model, unlike the fixed-sampling test.
The fixed-sampling test sends one prompt ~1000 times, builds a rarefaction curve of unique outputs, and Mann-Whitney tests it against the identity line.
Not vendored here — clone it to third_party/watermark-detection, or point
WATERMARK_DETECTION_DIR at it. Three gotchas found the hard way:
closed_models.ipynb imports src.attacks.kgw_detection and
stanford_detection; the real modules are redgreen_test and
fixedsampling_test. The notebook has not been run in a while.mangoes/pineapples/papayas/kiwis vs strawberries/cherries/mangoes/plums),
which silently yields all-zero counts for three of four words. Use one list.redgreen_test.py imports torch and the whole src.kgw package at module
level, none of which the statistic uses. src/their_stat.py extracts the two
functions by AST and execs them, so the bodies stay byte-identical without
needing a torch install.Their logit is log((p + 1e-10) / (1 - p + 1e-10)), not scipy.special.logit.
It saturates at ±23.03 instead of going infinite, so it tolerates proportions of
exactly 0 and 1 — but see findings.md for why that tolerance is
not the same as usefulness.
Jovanović, Staab, Vechev — ICML 2024. Same lab.
Approximates the watermark rules by querying the model, enabling both spoofing and scrubbing. For our purposes the interesting half is that it reconstructs an approximate detector from black-box access alone — which is precisely the goal here, without waiting for Anthropic’s documentation.
Critically, it operates on free-form text, where Claude has abundant entropy, rather than forced choice, where it does not. That sidesteps the obstacle that made the red-green test untestable on Opus 5. It is the natural next escalation; the ICML work used on the order of 2.2M tokens.
Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub