ccwatermark

Literature and methods

The three scheme families

Detection work classifies text watermarks into three families. Anthropic’s scheme is undisclosed, so all three are candidates — though press paraphrase (“subtly favour certain tokens”) points at the first.

Red-green / distribution-biasing (Kirchenbauer et al., ICML 2023)

Hash the last h tokens to seed a PRNG, split the vocabulary into a “green” list of fraction γ and a red remainder, add a bias δ to green logits. Detection counts green tokens against the binomial null. Parameters h, γ, δ are the scheme specification. Typical deployed δ ≈ 2.0.

Fixed-sampling (Aaronson; KTH)

Replace sampling randomness with a pseudorandom function of the context and a key. Output becomes deterministic given the context, so the same prompt yields the same completion. Distortion-free in expectation, but the determinism is a glaring tell.

Cache-augmented

Keep a cache of recently used watermark keys to avoid reuse. Detectable by comparing rapid queries (cache warm) against widely spaced ones (cache cold).

SynthID-Text (Google DeepMind, Nature 2024)

Tournament sampling. Deployed on Gemini, open-sourced, detector shipped in Hugging Face transformers. This is Google’s, not Anthropic’s — a common confusion in secondary coverage. Detection still requires the deployment’s keys.

Black-Box Detection of Language Model Watermarks

Gloaguen, Jovanović, Staab, Vechev — ICLR 2025. arXiv:2405.20777 · code

The core dependency of this repository. Statistical tests that detect the presence of a watermark in a deployed API using only sampled text — no logprobs, no key. Validated against seven schemes and five open models, then applied to GPT-4, Claude 3 and Gemini 1.0 Pro in 2024, i.e. before Anthropic deployed anything.

The red-green test is the clever one. Force a “random” pick from a four-word list while controlling the tokens immediately preceding the choice with a repeated-digit nonsense string (I ate 111111111111111, then 222..., through 999...). Because the string is a repeated digit, the last h tokens are the same whatever h is, so the green list depends only on which digit — the test is robust to an unknown context width. Each digit induces a different green list, so a given word is boosted under some digits and suppressed under others, consistently across all sentence prefixes. Without a watermark, the choice should not care which digit precedes it.

The statistic (test_kgw_detection) logit-transforms the per-context proportions, flags cells more than 1.96σ from the median, and generates its null by permuting the digit alignment within the observed data. That self- calibration is important: the red-green test needs no trusted unwatermarked control model, unlike the fixed-sampling test.

The fixed-sampling test sends one prompt ~1000 times, builds a rarefaction curve of unique outputs, and Mann-Whitney tests it against the identity line.

Working with their code

Not vendored here — clone it to third_party/watermark-detection, or point WATERMARK_DETECTION_DIR at it. Three gotchas found the hard way:

Their logit is log((p + 1e-10) / (1 - p + 1e-10)), not scipy.special.logit. It saturates at ±23.03 instead of going infinite, so it tolerates proportions of exactly 0 and 1 — but see findings.md for why that tolerance is not the same as usefulness.

Watermark Stealing in Large Language Models

Jovanović, Staab, Vechev — ICML 2024. Same lab.

Approximates the watermark rules by querying the model, enabling both spoofing and scrubbing. For our purposes the interesting half is that it reconstructs an approximate detector from black-box access alone — which is precisely the goal here, without waiting for Anthropic’s documentation.

Critically, it operates on free-form text, where Claude has abundant entropy, rather than forced choice, where it does not. That sidesteps the obstacle that made the red-green test untestable on Opus 5. It is the natural next escalation; the ICML work used on the order of 2.2M tokens.

Other relevant work


Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub