Everything collected so far is a null. Nulls accumulate but do not converge — the 2026-08-10 cycle can be summarised as “four ways of not finding it”. This cycle is designed to either find a mark or narrow the hypothesis space enough that the next null means something.
Three changes drive it:
The red-green statistic needs the same context sampled many times, so the distribution at that context can be estimated without logprobs. Long free-form text gives every token a unique context, so the estimator has nothing to average over. That is why the 2026-08-10 design used a 9-token forced choice — and why it landed in the exempt regime.
These pull in opposite directions only if the probe must be the whole output. It does not.
Status 2026-08-11, after running the probes. Design 1 below is dead — all four models round-robin the list instead of sampling it (0/152 repeats, p = 2e-19). It is kept here because the failure is instructive and the pre-registered check is what caught it. Design 1b replaced it and works on Haiku. See findings.md.
Have one call emit many controlled decision points inside one long output.
Write exactly 20 numbered lines. Each line has the form:
<n>. I ate <15 repeated digits> <fruit>.
Choose <fruit> from: mangoes, pineapples, papayas, kiwis.
Why this works:
h tokens, and h is small (typically 1–4). Fifteen repeated digits swamp
that window, so the effective context at the fruit token is the digit — exactly
as in the 2026-08-10 design, and independent of the line’s position.Watch for:
analyse_redgreen.py currently bootstraps over samples and would overstate
significance if pointed at this data unchanged.One probe per response, placed after ≥220 words of story, so there are no sibling
choices for the model to balance against. sample_longform.py --mode preamble.
./.venv/bin/python src/sample_longform.py --model claude-haiku-4-5 \
--mode preamble --digits 9 --calls 4500 --concurrency 12 \
--out data/$(date +%F)/lf_haiku45.jsonl
Verified on 120 live calls: median 331 output tokens on Haiku, 100% parse, 4/4 words live, max share 0.60–0.65 per digit. Better dynamic range than the original forced choice and above the 200-token threshold.
Two consequences worth noting:
analyse_redgreen.py needs no change — the clustering fix that
Design 1 required is off the critical path.The single most informative variable available, and it is nearly free once Design 1 works. Run the same probe at total output lengths spanning the threshold:
| Lines per call | Approx output tokens | Regime |
|---|---|---|
| 2 | ~25 | exempt |
| 8 | ~100 | exempt |
| 16 | ~200 | boundary |
| 30 | ~380 | mandated |
| 60 | ~750 | mandated, high reliability |
Then read the result off the shape:
Note the sweep is confounded with per-line position: line 2 of 60 sits in a different place than line 2 of 2. Analyse at matched line indices across lengths, not pooled.
The deepest weakness of every result so far is that the harness has never detected a real watermark — only synthetic ones it generated itself.
The original plan was to buy Gemini calls and hope. Reading the SynthID source produced something better and free: the reference implementation runs locally, GPT-2 on CPU. See synthid.md. That gives ground truth with known keys, watermarked and unwatermarked text from the same model, and no API budget.
It also produced a result that outranks everything else in this file:
Our repeated-digit context is skipped by SynthID by construction. The scheme does not watermark a token whose preceding 4-gram was recently seen, and a run of five identical digit tokens makes the fruit token’s context a repeat of the one immediately before it. The signal is absent at the exact position we measure.
Fix the stimulus before spending anything. Replace the repeated digits with nine fixed nonce phrases whose tokens are distinct. Same context control, no repetition for a guard to skip, and it strictly dominates the digit for KGW too.
Order: run the local control, quantify the digit-vs-nonce gap on ground truth, then collect.
Still unsolved after six attempts. Under the preamble design Opus 5 improved from
95% to 78% on papayas but left mangoes and pineapples never chosen — 2/4
words on a constant rail, which is still degenerate. Framing is no longer the
bottleneck; the word list is.
papayas prior in this frame; a list without a runaway favourite may clear 0.8.
Cheapest untried lever, ~$0.50 per candidate list.Revised 2026-08-11 after the probes and the SynthID read.
Nothing paid should happen before step 3 clears.
Stated in advance, so it cannot be chosen after seeing the data:
Anything less is a lead, and gets recorded as post-hoc — see the digit-1 anomaly in findings.md for how easily this project generates those.
Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub