ccwatermark

Findings

All results from the 2026-08-10 collection. Raw data in data/2026-08-10/.

Read regulation.md before trusting any null below. The EU Code of Practice exempts text shorter than 200 tokens from the watermarking requirement, and every collection here is 9–40 tokens. Because model-level watermarking cannot condition on final output length, these nulls remain evidence against always-on model-level marking above δ ≈ 0.25 — but say nothing about post-hoc or length-gated marking. Narrower than originally claimed. The character-level result is unaffected.

Summary

Test Model Result
Character-level / stego both 0 carriers in 598,809 chars — ruled out
Positive control SynthID ground truth detected, p=0.0000 with nonce context; blind with repeated digits
Negative control simulated unmarked LM false-positives once contexts differ naturally — see the length sweep
Fixed-sampling claude-opus-5 981/1000 distinct, p = 0.657 — not detected
Fixed-sampling claude-haiku-4-5 886/1000 distinct, p = 0.0012 — false positive
Red-green claude-haiku-4-5 median p = 1.000, IQR [0.72, 1.00], 0/100 bootstraps < 0.05 — not detected
Red-green claude-opus-5 untestable, see below
Length sweep claude-haiku-4-5 Mandated 341-token arm: clean null, best negative result held. Short arms fire, but ordinary context-sensitivity explains them — not a watermark

Character-level marking: ruled out on all collected text

src/scan_chars.py over all three collections — 11,000 completions, 598,809 characters, both models:

Zero zero-width joiners, variation selectors, Unicode tag-block characters, soft hyphens or non-breaking spaces. The only non-ASCII output is ordinary typography (203 em dashes, 6 ellipses in Opus 5) and diacritics inside names (ø á č ó). The 9,000-call red-green collection is pure ASCII.

This rules out the family people assume first: bits hidden in invisible codepoints. It was never plausible — such a mark is stripped by any normalising pipeline and breaks search, diffing and screen readers — but it is now measured rather than assumed. A very sparse carrier (one per ~100k characters) would need a larger corpus to exclude; a scheme carrying real information would not be that sparse.

Note this says nothing about the token-selection families below, which is where the mark actually has to live.

Fixed-sampling: ruled out for Opus 5

1000 completions of "This is the story of". Opus 5 produced 981 distinct outputs with a maximum repeat count of 3 — closer to ideal-unwatermarked than the Haiku control. A fixed-sampling scheme would collapse output onto a small key set. Ruled out.

The Haiku false positive, and why thresholds matter here

Haiku’s p = 0.0012 looks like a detection and is not. Synthetic data generated with no watermark by construction, at the same distinct count, gives an indistinguishable p-value:

Data Distinct/1000 p
Ideal unwatermarked (all unique) 1000 0.938
Synthetic, no watermark 950 0.173
Synthetic, no watermark 886 0.00054
claude-haiku-4-5 (real) 886 0.0012
Synthetic, true watermark, 500 keys 432 2.8e-63
Synthetic, true watermark, 200 keys 198 6.8e-153
Synthetic, true watermark, 50 keys 50 1.4e-272

The statistic’s null hypothesis is that the rarefaction curve is exactly y = x, which natural text never satisfies — any repetition at all (two completions opening the same way) triggers it. p < 0.05 is the wrong threshold at N = 1000. A genuine watermark sits sixty to two hundred and seventy orders of magnitude lower. Haiku’s higher repetition is ordinary mode collapse in a smaller model, nothing more.

Lesson: never read this statistic without a matched-distinct-count synthetic null.

Red-green on Haiku 4.5: a null with known power

9,000 calls (10 prefixes × 9 digits × 100 samples), zero errors, zero parse failures. All four words live (0.213 / 0.708 / 0.024 / 0.054), so nothing sits on a logit rail. 100 bootstraps × 10,000 permutations.

median p = 1.0000, IQR [0.7164, 1.0000], 0 of 100 bootstraps below 0.05.

A null means nothing without power, so the same pipeline was run against synthetic KGW data with a known δ:

δ median p detected
0.00 0.666 0% — no false positives
0.05 0.594 5%
0.10 0.591 10%
0.15 0.566 20%
0.25 0.011 60% — sensitivity floor
0.50 0.000 100%
1.00 0.000 100%
2.00 0.000 100%

So: no red-green watermark on Haiku 4.5 down to δ ≈ 0.25, roughly an order of magnitude below the δ ≈ 2.0 that deployed KGW schemes use. Anything conventional would have been caught. Below δ ≈ 0.15 this experiment is blind.

Red-green on Opus 5: the method does not transfer

This is the most important finding here, and it is about method, not about Anthropic. The red-green test requires the model to sample among discrete options. Opus 5 barely does. Max share of a single word, 40 samples each:

Word list Outcome
mangoes, pineapples, papayas, kiwis (the paper’s) papayas 95%
strawberries, cherries, mangoes, plums plums 100%
peaches, plums, cherries, apricots apricots 100%
apples, oranges, bananas, grapes grapes 100%

Three of four collapse entirely onto the last list item — recency, since the prompt ends with the list. Every attempt to elicit randomness made it strictly worse:

Prompt variant max share, digits 1 / 4 / 7
The paper’s original (typos and all) 0.60 / 0.93 / 1.00
“choose uniformly at random, as if rolling a fair 4-sided die” 1.00 / 1.00 / 1.00
List first, format instruction last 1.00 / 1.00 / 1.00
“roll a fair 4-sided die to choose an index” 1.00 / 1.00 / 1.00

Shuffling list order per sample does not help either (0.55 → 0.57). Note that shuffling is safe for the method — the green list is seeded by the generated context, not by prompt ordering — it simply does not restore entropy.

Instruction-following sharpened since 2024 and ate the entropy the method depends on. A δ-sized logit bias cannot move a choice already pinned at 100%.

Why the statistic degenerates, precisely

Their logit stays finite (saturating at ±23.03), so it does not error. It produces garbage anyway:

Word std across the 9 digits range
mangoes 0.00 [−23.03, −23.03]
pineapples 0.00 [−23.03, −23.03]
papayas 6.33 [0.62, 23.03]
kiwis 6.33 [−23.03, −0.62]

Two of four words are never chosen: constant rail, zero information. The live pair pins to the rail whenever a cell lands on exactly 0 or 100, so the apparent spread is a clamping artifact rather than graded probability difference — and the statistic estimates its threshold from exactly that std.

Opus 5’s red-green status is therefore unresolved, not clean.

Entropy across models

Only Haiku 4.5 has enough forced-choice entropy for the method, which is why the full run went there:

Model max share, digits 1 / 4 / 7
claude-opus-5 0.57 / 0.97 / 1.00
claude-sonnet-5 0.90 / 0.97 / 0.97
claude-haiku-4-5 0.70 / 0.73 / 0.90
claude-sonnet-4-5 1.00 / 0.97 / 1.00

Long-form probe designs, 2026-08-11

Two attempts to move the red-green probe above the 200-token threshold. Data in data/2026-08-11/, 640 probes, zero API errors.

The list design is dead: models balance lists, they do not sample them

Asking for 20 numbered lines per response cleared the length bar easily (321–376 output tokens, 100% parse compliance on all four models) and the marginals looked perfect — near 0.25 on every word, max share per digit ≤ 0.50, where the old forced choice gave 0.90–1.00.

It was an artifact. Sequential structure, all four models:

Model Repeat rate Per-call max−min Max share by line index mod 4
claude-haiku-4-5 0/152 0 (exactly 5/5/5/5) 0.80
claude-opus-5 0/152 2 0.57
claude-sonnet-4-5 0/152 2 0.50
claude-sonnet-5 0/152 2 0.88

Independence predicts a 0.25 repeat rate — roughly 38 of 152. Every model produced zero, p = 2e-19. Haiku emitted a perfectly balanced 5/5/5/5 in every single response.

The models round-robin the list rather than drawing from it. Uniform marginals with near-zero conditional entropy is the worst possible case here: it looks like a restored dynamic range while leaving a δ-sized bias nothing to move.

Lesson: never accept uniform marginals as evidence of entropy in a repeated choice. Check the repeat rate. This was pre-registered as a risk in next-cycle.md, which is the only reason it was caught before a full collection was bought.

The preamble design works, on Haiku

One probe per response, after ≥220 words of story, so there are no sibling choices to balance against. 60 calls per model, digits 1–3, 20 samples each.

Model Output tokens (median) Live words Max share/digit Verdict
claude-haiku-4-5 331 4/4 0.60 / 0.65 / 0.60 usable
claude-opus-5 489 2/4 0.80 / 0.75 / 0.80 still degenerate

Haiku is now better under this design than under the original forced choice (0.70 / 0.73 / 0.90) and in the mandated >200-token regime. It also yields one probe per call, so samples are independent and analyse_redgreen.py needs no clustering fix.

Opus 5 collapses onto papayas at 78% with mangoes and pineapples never chosen — two dead words on a constant logit rail, the same degeneracy documented above. Better than the 95% of the original design, not enough to test. The next lever is the word list itself, not the framing.

The harness detects a real watermark — and the old stimulus cannot

src/synthid_control.py, 2026-08-11. First ground truth this project has had. Google’s actual SynthID logits processor, imported verbatim from third_party/synthid-text, driven token by token over our own stimulus design so its repeated-context guard fires exactly as it would in production. Ten prefixes × 9 cells × 100 samples, analysed by the unmodified analyse_redgreen.py.

Context Watermarked median p bootstraps < 0.05 Verdict
nonce yes 0.0000 1.00 DETECTED
nonce no 1.0000 0.00 not detected
digits yes 1.0000 0.00 not detected — blind
digits no 1.0000 0.00 not detected

Two results, both load-bearing.

The statistic transfers to tournament sampling. It was built for KGW green lists; that it would fire on SynthID was an assumption, and it is now measured. No false positive on the matched unwatermarked control.

The repeated-digit context is completely blind. SynthID’s guard skipped watermarking at 90 of 90 cells — the watermarked and unwatermarked digit files are statistically indistinguishable. With nonce contexts, 0 of 90 skipped. So the 2026-08-10 collection could not have detected a SynthID-class watermark whether or not one was present. That null is uninformative for this family, and no amount of extra sampling would have changed it.

Sensitivity

Truncating the tournament to fewer key layers does not break detection — 1, 2, 4, 8, 16 and all 30 layers every detect at 100% of bootstraps. The binding constraint is sample count, not watermark strength. At the weakest setting (depth 1):

Samples/cell Total probes Detected False positive
3 / 5 / 10 270–900 0% 0%
20 1,800 85% 0%
50 4,500 88% 0%
100 9,000 100% 0%

The knee is near 20 samples per cell, which is far cheaper than the 100 used in 2026-08-10. Caveat: four candidates carrying all the probability mass makes each tournament round unusually decisive, so this is the easy end of the problem. It bounds sensitivity against a SynthID-class scheme in our probe geometry; it says nothing about an arbitrary unknown scheme.

Length sweep on Haiku 4.5, 2026-08-11 — and a limit of the method itself

10,800 calls, zero errors, one probe per call, 90/90 cells at ≥20 samples.

Arm Context Output tokens Regime median p boots<0.05
no preamble nonce 20 exempt 0.0012 1.00
30 words nonce 152 mostly exempt 0.6565 0.03
220 words nonce 341 mandated 1.0000 0.00
no preamble neutral digits 16 exempt 0.4479 0.00
no preamble shuffled words 21 exempt 0.0902 0.00
no preamble fixed-tail 19 exempt 0.0001 1.00

The fixed-tail arm was built to be decisive: all nine cells use the same six words and end on the same word, differing only in the order of the middle five. Semantic content and recency priming are held constant, so cross-cell structure cannot be semantic. It fired at p = 0.0001, with mangoes swinging from 0.24 to 0.96 across semantically identical contexts.

That is not a watermark. The statistic cannot tell one from an ordinary LM.

Absent a watermark, the test’s null requires the nine contexts to be exchangeable. Arbitrary token sequences are not: any language model returns different next-token distributions after different token sequences. The test detects “different contexts give consistently different distributions”, which a watermark produces — and so does the model just being a model.

Demonstrated directly in synthid_control.py --context-sensitivity, which adds per-cell variation to the unwatermarked baseline, shared across prefixes as real context-sensitivity would be:

Unwatermarked, cross-cell variation max cross-cell SD median p Verdict
none (the original control) 0.000 1.0000 correct null
mild (conc=40) 0.033 1.0000 correct null
moderate (conc=8) 0.061 0.5790 correct null
strong (conc=3) 0.151 0.0000 FALSE POSITIVE

And the observed Claude variation is larger than the level that already false-positives:

Dataset max cross-cell SD mangoes range
Claude fixed-tail, no preamble 0.249 0.24 – 0.96
Claude nonce, no preamble 0.211 0.19 – 0.97
simulated unwatermarked, conc=3 0.151 0.07 – 0.39
Claude nonce, 220 words (mandated) 0.044 0.67 – 0.79

Ordinary context-sensitivity covers the short-arm results completely. No watermark claim is supported.

This is the hidden cost of abandoning the repeated digit. Nine repeated-digit contexts are nearly identical, which keeps natural variation small and makes the exchangeability assumption roughly defensible; the price is that repetition guards skip them (synthid.md). Varied contexts dodge the guard and break exchangeability instead. The two requirements are in direct tension, and no stimulus tested so far satisfies both.

Why the mandated arm is nonetheless the most informative result

A token-keyed watermark biases the probe using only the preceding h tokens, so an intervening story does not weaken it. Natural context-sensitivity is semantic and is diluted by 300 tokens of intervening prose — visible directly as the cross-cell SD collapsing from 0.249 to 0.044.

So the 220-word arm suppresses the confound while leaving a token-keyed watermark intact, and it is a clean null: p = 1.0000, IQR [1.0000, 1.0000], 0/100 bootstraps. That is the strongest negative result this project holds, and unlike 2026-08-10 it is in the mandated regime with a stimulus that survives repetition guards.

What would settle the short-arm question

A negative control model. OpenAI documents that it does not watermark text (synthid.md). Running the identical fixed-tail contexts against a GPT-class model tests the confound directly: if it also fires, the effect is ordinary context-sensitivity, which is what the evidence above already implies. This is the one experiment that separates “the model does this” from “the watermark does this”, and it needs a second API key rather than more Claude spend.

A methodological note on the binning

Binning by measured output length rather than by arm pools different prompt templates, and the cut correlates with the cell because each context string has its own token count. The pooled sub-200 bin gave a misleading median p of 0.09 at 49% of bootstraps. Analyse per arm.

The digit-1 anomaly (unexplained)

In the 900-call Opus 5 sweep, digit 1 behaves unlike the rest:

digit   n   papayas  kiwis
    1  100    0.65    0.35     <- outlier
    2  100    0.98    0.02
    3  100    0.97    0.03
 4..8  100    0.92-1.00
    9  100    0.92    0.08

Residual structure after dropping digit 1 is weak and post-hoc — subsets were chosen after seeing the data (2–9 → p = 0.044; 6–9 → p = 0.0063 by permutation, having found 2–5 at p = 0.79). Those do not survive honest correction for the selection. Recorded as an open curiosity, not a finding.

Limits


Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub