All results from the 2026-08-10 collection. Raw data in data/2026-08-10/.
Read regulation.md before trusting any null below. The EU Code of Practice exempts text shorter than 200 tokens from the watermarking requirement, and every collection here is 9–40 tokens. Because model-level watermarking cannot condition on final output length, these nulls remain evidence against always-on model-level marking above δ ≈ 0.25 — but say nothing about post-hoc or length-gated marking. Narrower than originally claimed. The character-level result is unaffected.
| Test | Model | Result |
|---|---|---|
| Character-level / stego | both | 0 carriers in 598,809 chars — ruled out |
| Positive control | SynthID ground truth | detected, p=0.0000 with nonce context; blind with repeated digits |
| Negative control | simulated unmarked LM | false-positives once contexts differ naturally — see the length sweep |
| Fixed-sampling | claude-opus-5 |
981/1000 distinct, p = 0.657 — not detected |
| Fixed-sampling | claude-haiku-4-5 |
886/1000 distinct, p = 0.0012 — false positive |
| Red-green | claude-haiku-4-5 |
median p = 1.000, IQR [0.72, 1.00], 0/100 bootstraps < 0.05 — not detected |
| Red-green | claude-opus-5 |
untestable, see below |
| Length sweep | claude-haiku-4-5 |
Mandated 341-token arm: clean null, best negative result held. Short arms fire, but ordinary context-sensitivity explains them — not a watermark |
src/scan_chars.py over all three collections — 11,000 completions, 598,809
characters, both models:
Zero zero-width joiners, variation selectors, Unicode tag-block characters,
soft hyphens or non-breaking spaces. The only non-ASCII output is ordinary
typography (203 em dashes, 6 ellipses in Opus 5) and diacritics inside names
(ø á č ó). The 9,000-call red-green collection is pure ASCII.
This rules out the family people assume first: bits hidden in invisible codepoints. It was never plausible — such a mark is stripped by any normalising pipeline and breaks search, diffing and screen readers — but it is now measured rather than assumed. A very sparse carrier (one per ~100k characters) would need a larger corpus to exclude; a scheme carrying real information would not be that sparse.
Note this says nothing about the token-selection families below, which is where the mark actually has to live.
1000 completions of "This is the story of". Opus 5 produced 981 distinct
outputs with a maximum repeat count of 3 — closer to ideal-unwatermarked than the
Haiku control. A fixed-sampling scheme would collapse output onto a small key set.
Ruled out.
Haiku’s p = 0.0012 looks like a detection and is not. Synthetic data generated with no watermark by construction, at the same distinct count, gives an indistinguishable p-value:
| Data | Distinct/1000 | p |
|---|---|---|
| Ideal unwatermarked (all unique) | 1000 | 0.938 |
| Synthetic, no watermark | 950 | 0.173 |
| Synthetic, no watermark | 886 | 0.00054 |
claude-haiku-4-5 (real) |
886 | 0.0012 |
| Synthetic, true watermark, 500 keys | 432 | 2.8e-63 |
| Synthetic, true watermark, 200 keys | 198 | 6.8e-153 |
| Synthetic, true watermark, 50 keys | 50 | 1.4e-272 |
The statistic’s null hypothesis is that the rarefaction curve is exactly y = x,
which natural text never satisfies — any repetition at all (two completions
opening the same way) triggers it. p < 0.05 is the wrong threshold at
N = 1000. A genuine watermark sits sixty to two hundred and seventy orders of
magnitude lower. Haiku’s higher repetition is ordinary mode collapse in a smaller
model, nothing more.
Lesson: never read this statistic without a matched-distinct-count synthetic null.
9,000 calls (10 prefixes × 9 digits × 100 samples), zero errors, zero parse failures. All four words live (0.213 / 0.708 / 0.024 / 0.054), so nothing sits on a logit rail. 100 bootstraps × 10,000 permutations.
median p = 1.0000, IQR [0.7164, 1.0000], 0 of 100 bootstraps below 0.05.
A null means nothing without power, so the same pipeline was run against
synthetic KGW data with a known δ:
| δ | median p | detected |
|---|---|---|
| 0.00 | 0.666 | 0% — no false positives |
| 0.05 | 0.594 | 5% |
| 0.10 | 0.591 | 10% |
| 0.15 | 0.566 | 20% |
| 0.25 | 0.011 | 60% — sensitivity floor |
| 0.50 | 0.000 | 100% |
| 1.00 | 0.000 | 100% |
| 2.00 | 0.000 | 100% |
So: no red-green watermark on Haiku 4.5 down to δ ≈ 0.25, roughly an order of magnitude below the δ ≈ 2.0 that deployed KGW schemes use. Anything conventional would have been caught. Below δ ≈ 0.15 this experiment is blind.
This is the most important finding here, and it is about method, not about Anthropic. The red-green test requires the model to sample among discrete options. Opus 5 barely does. Max share of a single word, 40 samples each:
| Word list | Outcome |
|---|---|
mangoes, pineapples, papayas, kiwis (the paper’s) |
papayas 95% |
strawberries, cherries, mangoes, plums |
plums 100% |
peaches, plums, cherries, apricots |
apricots 100% |
apples, oranges, bananas, grapes |
grapes 100% |
Three of four collapse entirely onto the last list item — recency, since the prompt ends with the list. Every attempt to elicit randomness made it strictly worse:
| Prompt variant | max share, digits 1 / 4 / 7 |
|---|---|
| The paper’s original (typos and all) | 0.60 / 0.93 / 1.00 |
| “choose uniformly at random, as if rolling a fair 4-sided die” | 1.00 / 1.00 / 1.00 |
| List first, format instruction last | 1.00 / 1.00 / 1.00 |
| “roll a fair 4-sided die to choose an index” | 1.00 / 1.00 / 1.00 |
Shuffling list order per sample does not help either (0.55 → 0.57). Note that shuffling is safe for the method — the green list is seeded by the generated context, not by prompt ordering — it simply does not restore entropy.
Instruction-following sharpened since 2024 and ate the entropy the method depends on. A δ-sized logit bias cannot move a choice already pinned at 100%.
Their logit stays finite (saturating at ±23.03), so it does not error. It
produces garbage anyway:
| Word | std across the 9 digits | range |
|---|---|---|
| mangoes | 0.00 | [−23.03, −23.03] |
| pineapples | 0.00 | [−23.03, −23.03] |
| papayas | 6.33 | [0.62, 23.03] |
| kiwis | 6.33 | [−23.03, −0.62] |
Two of four words are never chosen: constant rail, zero information. The live pair pins to the rail whenever a cell lands on exactly 0 or 100, so the apparent spread is a clamping artifact rather than graded probability difference — and the statistic estimates its threshold from exactly that std.
Opus 5’s red-green status is therefore unresolved, not clean.
Only Haiku 4.5 has enough forced-choice entropy for the method, which is why the full run went there:
| Model | max share, digits 1 / 4 / 7 |
|---|---|
claude-opus-5 |
0.57 / 0.97 / 1.00 |
claude-sonnet-5 |
0.90 / 0.97 / 0.97 |
claude-haiku-4-5 |
0.70 / 0.73 / 0.90 |
claude-sonnet-4-5 |
1.00 / 0.97 / 1.00 |
Two attempts to move the red-green probe above the 200-token threshold. Data in
data/2026-08-11/, 640 probes, zero API errors.
Asking for 20 numbered lines per response cleared the length bar easily (321–376 output tokens, 100% parse compliance on all four models) and the marginals looked perfect — near 0.25 on every word, max share per digit ≤ 0.50, where the old forced choice gave 0.90–1.00.
It was an artifact. Sequential structure, all four models:
| Model | Repeat rate | Per-call max−min | Max share by line index mod 4 |
|---|---|---|---|
claude-haiku-4-5 |
0/152 | 0 (exactly 5/5/5/5) | 0.80 |
claude-opus-5 |
0/152 | 2 | 0.57 |
claude-sonnet-4-5 |
0/152 | 2 | 0.50 |
claude-sonnet-5 |
0/152 | 2 | 0.88 |
Independence predicts a 0.25 repeat rate — roughly 38 of 152. Every model produced zero, p = 2e-19. Haiku emitted a perfectly balanced 5/5/5/5 in every single response.
The models round-robin the list rather than drawing from it. Uniform marginals with near-zero conditional entropy is the worst possible case here: it looks like a restored dynamic range while leaving a δ-sized bias nothing to move.
Lesson: never accept uniform marginals as evidence of entropy in a repeated choice. Check the repeat rate. This was pre-registered as a risk in next-cycle.md, which is the only reason it was caught before a full collection was bought.
One probe per response, after ≥220 words of story, so there are no sibling choices to balance against. 60 calls per model, digits 1–3, 20 samples each.
| Model | Output tokens (median) | Live words | Max share/digit | Verdict |
|---|---|---|---|---|
claude-haiku-4-5 |
331 | 4/4 | 0.60 / 0.65 / 0.60 | usable |
claude-opus-5 |
489 | 2/4 | 0.80 / 0.75 / 0.80 | still degenerate |
Haiku is now better under this design than under the original forced choice
(0.70 / 0.73 / 0.90) and in the mandated >200-token regime. It also yields
one probe per call, so samples are independent and analyse_redgreen.py needs no
clustering fix.
Opus 5 collapses onto papayas at 78% with mangoes and pineapples never
chosen — two dead words on a constant logit rail, the same degeneracy documented
above. Better than the 95% of the original design, not enough to test. The next
lever is the word list itself, not the framing.
src/synthid_control.py, 2026-08-11. First ground truth this project has had.
Google’s actual SynthID logits processor, imported verbatim from
third_party/synthid-text, driven token by token over our own stimulus design so
its repeated-context guard fires exactly as it would in production. Ten prefixes
× 9 cells × 100 samples, analysed by the unmodified analyse_redgreen.py.
| Context | Watermarked | median p | bootstraps < 0.05 | Verdict |
|---|---|---|---|---|
| nonce | yes | 0.0000 | 1.00 | DETECTED |
| nonce | no | 1.0000 | 0.00 | not detected |
| digits | yes | 1.0000 | 0.00 | not detected — blind |
| digits | no | 1.0000 | 0.00 | not detected |
Two results, both load-bearing.
The statistic transfers to tournament sampling. It was built for KGW green lists; that it would fire on SynthID was an assumption, and it is now measured. No false positive on the matched unwatermarked control.
The repeated-digit context is completely blind. SynthID’s guard skipped watermarking at 90 of 90 cells — the watermarked and unwatermarked digit files are statistically indistinguishable. With nonce contexts, 0 of 90 skipped. So the 2026-08-10 collection could not have detected a SynthID-class watermark whether or not one was present. That null is uninformative for this family, and no amount of extra sampling would have changed it.
Truncating the tournament to fewer key layers does not break detection — 1, 2, 4, 8, 16 and all 30 layers every detect at 100% of bootstraps. The binding constraint is sample count, not watermark strength. At the weakest setting (depth 1):
| Samples/cell | Total probes | Detected | False positive |
|---|---|---|---|
| 3 / 5 / 10 | 270–900 | 0% | 0% |
| 20 | 1,800 | 85% | 0% |
| 50 | 4,500 | 88% | 0% |
| 100 | 9,000 | 100% | 0% |
The knee is near 20 samples per cell, which is far cheaper than the 100 used in 2026-08-10. Caveat: four candidates carrying all the probability mass makes each tournament round unusually decisive, so this is the easy end of the problem. It bounds sensitivity against a SynthID-class scheme in our probe geometry; it says nothing about an arbitrary unknown scheme.
10,800 calls, zero errors, one probe per call, 90/90 cells at ≥20 samples.
| Arm | Context | Output tokens | Regime | median p | boots<0.05 |
|---|---|---|---|---|---|
| no preamble | nonce | 20 | exempt | 0.0012 | 1.00 |
| 30 words | nonce | 152 | mostly exempt | 0.6565 | 0.03 |
| 220 words | nonce | 341 | mandated | 1.0000 | 0.00 |
| no preamble | neutral digits | 16 | exempt | 0.4479 | 0.00 |
| no preamble | shuffled words | 21 | exempt | 0.0902 | 0.00 |
| no preamble | fixed-tail | 19 | exempt | 0.0001 | 1.00 |
The fixed-tail arm was built to be decisive: all nine cells use the same six
words and end on the same word, differing only in the order of the middle
five. Semantic content and recency priming are held constant, so cross-cell
structure cannot be semantic. It fired at p = 0.0001, with mangoes swinging
from 0.24 to 0.96 across semantically identical contexts.
Absent a watermark, the test’s null requires the nine contexts to be exchangeable. Arbitrary token sequences are not: any language model returns different next-token distributions after different token sequences. The test detects “different contexts give consistently different distributions”, which a watermark produces — and so does the model just being a model.
Demonstrated directly in synthid_control.py --context-sensitivity, which adds
per-cell variation to the unwatermarked baseline, shared across prefixes as
real context-sensitivity would be:
| Unwatermarked, cross-cell variation | max cross-cell SD | median p | Verdict |
|---|---|---|---|
| none (the original control) | 0.000 | 1.0000 | correct null |
| mild (conc=40) | 0.033 | 1.0000 | correct null |
| moderate (conc=8) | 0.061 | 0.5790 | correct null |
| strong (conc=3) | 0.151 | 0.0000 | FALSE POSITIVE |
And the observed Claude variation is larger than the level that already false-positives:
| Dataset | max cross-cell SD | mangoes range |
|---|---|---|
| Claude fixed-tail, no preamble | 0.249 | 0.24 – 0.96 |
| Claude nonce, no preamble | 0.211 | 0.19 – 0.97 |
| simulated unwatermarked, conc=3 | 0.151 | 0.07 – 0.39 |
| Claude nonce, 220 words (mandated) | 0.044 | 0.67 – 0.79 |
Ordinary context-sensitivity covers the short-arm results completely. No watermark claim is supported.
This is the hidden cost of abandoning the repeated digit. Nine repeated-digit contexts are nearly identical, which keeps natural variation small and makes the exchangeability assumption roughly defensible; the price is that repetition guards skip them (synthid.md). Varied contexts dodge the guard and break exchangeability instead. The two requirements are in direct tension, and no stimulus tested so far satisfies both.
A token-keyed watermark biases the probe using only the preceding h tokens, so
an intervening story does not weaken it. Natural context-sensitivity is
semantic and is diluted by 300 tokens of intervening prose — visible directly
as the cross-cell SD collapsing from 0.249 to 0.044.
So the 220-word arm suppresses the confound while leaving a token-keyed watermark intact, and it is a clean null: p = 1.0000, IQR [1.0000, 1.0000], 0/100 bootstraps. That is the strongest negative result this project holds, and unlike 2026-08-10 it is in the mandated regime with a stimulus that survives repetition guards.
A negative control model. OpenAI documents that it does not watermark text (synthid.md). Running the identical fixed-tail contexts against a GPT-class model tests the confound directly: if it also fires, the effect is ordinary context-sensitivity, which is what the evidence above already implies. This is the one experiment that separates “the model does this” from “the watermark does this”, and it needs a second API key rather than more Claude spend.
Binning by measured output length rather than by arm pools different prompt templates, and the cut correlates with the cell because each context string has its own token count. The pooled sub-200 bin gave a misleading median p of 0.09 at 49% of bootstraps. Analyse per arm.
In the 900-call Opus 5 sweep, digit 1 behaves unlike the rest:
digit n papayas kiwis
1 100 0.65 0.35 <- outlier
2 100 0.98 0.02
3 100 0.97 0.03
4..8 100 0.92-1.00
9 100 0.92 0.08
Residual structure after dropping digit 1 is weak and post-hoc — subsets were chosen after seeing the data (2–9 → p = 0.044; 6–9 → p = 0.0063 by permutation, having found 2–5 at p = 0.79). Those do not survive honest correction for the selection. Recorded as an open curiosity, not a finding.
MAX_TOKENS = 40; the red-green stimulus is ~9 tokens. Marking is not required
below 200. See regulation.md.Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub