The ETH code targets a 2024 API. Three things changed and all three matter.
temperature is rejected. Removed on Opus 5, Sonnet 5, Opus 4.7+ and Fable 5
— sending it returns a 400. The published code passes temperature=1.0
everywhere. Omit it; the default is stochastic, which the fixed-sampling result
confirms empirically (981/1000 distinct).
No logprobs, ever. Anthropic has never exposed them. Every distribution must be estimated by repeated sampling, which sets the sample counts here.
Thinking must be explicitly disabled. On Opus 5 and Sonnet 5 thinking is on by default; its tokens would sit between the controlled context and the sampled token, destroying any test that depends on the preceding tokens. Per model:
| Model | Setting |
|---|---|
claude-opus-5, claude-sonnet-5, claude-opus-4-8 |
thinking={"type": "disabled"} |
claude-haiku-4-5, claude-sonnet-4-5, claude-opus-4-5 |
omit — these do not think unless asked |
On Opus 5, disabled thinking is accepted only at effort high or below; the
default effort is high, so passing nothing else works. Pairing it with xhigh
or max returns a 400.
uv venv --python 3.12 .venv
uv pip install --python ./.venv/bin/python -r requirements.txt
git clone https://github.com/eth-sri/watermark-detection third_party/watermark-detection
An env file holding ANTHROPIC_API_KEY exists outside this repo. Source it into
the environment before running; do not commit it, echo it, or paste it into a
file inside the repo.
set -a && . /path/to/your/env/file && set +a
src/their_stat.py finds the cloned repo at third_party/watermark-detection,
overridable with WATERMARK_DETECTION_DIR.
./.venv/bin/python src/sample_fixed.py --model claude-opus-5 --n 1000 \
--concurrency 8 --out data/$(date +%F)/fixed_opus5.jsonl
./.venv/bin/python src/analyse_fixed.py data/$(date +%F)/fixed_opus5.jsonl
Resumable — re-running counts existing lines and collects only the shortfall. ~1000 calls, minutes, about $1 on Opus 5.
Always analyse against a matched synthetic null. p < 0.05 alone is
meaningless here; see findings.md.
./.venv/bin/python src/sample_redgreen.py --model claude-haiku-4-5 \
--samples 100 --prefixes 10 --digits 9 --concurrency 12 \
--out data/$(date +%F)/rg_full_haiku45.jsonl
./.venv/bin/python src/analyse_redgreen.py data/$(date +%F)/rg_full_haiku45.jsonl \
--bootstrap 100 --permutations 10000
9,000 calls, ~25 minutes at concurrency 12, ~$2 on Haiku 4.5. The analysis is CPU-bound and takes a few minutes.
Check entropy before spending. A cheap probe (1 prefix, 3 digits, 30 samples) tells you whether the model is usable. If any word exceeds ~0.8 share the test has no dynamic range and the run is wasted — this is what makes Opus 5 untestable.
./.venv/bin/python src/synth_redgreen.py # writes synth_rg_delta*.jsonl
./.venv/bin/python src/analyse_redgreen.py synth_rg_delta0.0.jsonl \
--bootstrap 20 --permutations 2000
δ = 0.0 must come back null and δ ≥ 0.5 must come back detected. If not, the pipeline is broken and any null on real data is uninterpretable.
Route 1. Marking status changes over time, so re-running an identical protocol and differencing against an earlier dated collection detects activation with no disclosure from Anthropic.
./.venv/bin/python src/diff_redgreen.py \
data/2026-08-10/rg_full_haiku45.jsonl data/$(date +%F)/rg_full_haiku45.jsonl
Comparing the two runs’ detection p-values by eye does not work: both can sit at p = 1.0 while the proportions move, and a p-value has no sampling distribution you can subtract. The script tests the proportions directly, in three families:
| Family | Tables | n per side | Role |
|---|---|---|---|
| per-cell | 90, one per (prefix, digit) | ~100 | finest split, BH-corrected |
| per-digit | 9, prefixes pooled | ~1000 | where an activation shows first |
| omnibus | 1, joint over the summed G | — | the verdict |
Every table is tested by exact permutation of the collection labels under the fixed pooled margin (multivariate hypergeometric), never by a chi-square approximation — papayas sits near 2%, so expected counts fall below 5 and the approximation is invalid.
The per-digit family is the sensitive one, by design. A KGW green list seeded by the digit context lifts the same words at a given digit across every prefix, so pooling prefixes multiplies n tenfold without diluting the signal. At δ = 0.25 the per-digit route fires at p = 0.0008 on synthetic data where no single cell reaches uncorrected significance.
Run with one file to split it in half and diff it against itself. The null is true by construction, so this calibrates the false-positive rate on real data — the 2026-08-10 baseline gives 3/90 cells at uncorrected p < 0.05 against ~4 expected, omnibus p = 0.73.
A shift is not a watermark. A model update or a decoding change moves these numbers too. What separates them is shape: a green list lifts the same words at the same digit across all prefixes. When the verdict fires, the script prints a per-digit direction read-out with a sign test over prefixes — that consistency, not the p-value, is the evidence. On synthetic δ = 0.5 data it recovers the planted lists exactly (papayas 10/10 prefixes at digits 1/3/5/8, kiwis 10/10 at 2/7).
The verdict rule is an OR of two correlated tests, so its true α is not 0.05. Measured on the baseline it is ~7%. The δ = 0 row of the power curve reports this directly every run rather than assuming the nominal rate.
Never diff two synthetic files generated with the same seed. synth_redgreen.py
defaults to seed=0, so two files share a uniform stream and their draws are
correlated, not independent. The diff then reads more similar than chance —
omnibus p = 1 with 0/90 cells at uncorrected p < 0.05, below the ~4 expected.
Pass distinct seeds to gen() when constructing a test pair.
Power is reported at the full baseline n, even in split-half mode, because the weekly diff is 100-vs-100 and a curve measured at n = 50 would understate it.
Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub