ccwatermark

Next cycle: a design aimed at a positive result

Everything collected so far is a null. Nulls accumulate but do not converge — the 2026-08-10 cycle can be summarised as “four ways of not finding it”. This cycle is designed to either find a mark or narrow the hypothesis space enough that the next null means something.

Three changes drive it:

  1. Get above 200 tokens without giving up the controlled contexts the statistic needs.
  2. Discriminate between implementations rather than testing one family, using output length as the independent variable.
  3. Validate the harness against a watermark known to exist, so that a null on Claude is attributable to Claude rather than to our pipeline.

The core obstacle, stated precisely

The red-green statistic needs the same context sampled many times, so the distribution at that context can be estimated without logprobs. Long free-form text gives every token a unique context, so the estimator has nothing to average over. That is why the 2026-08-10 design used a 9-token forced choice — and why it landed in the exempt regime.

These pull in opposite directions only if the probe must be the whole output. It does not.

Status 2026-08-11, after running the probes. Design 1 below is dead — all four models round-robin the list instead of sampling it (0/152 repeats, p = 2e-19). It is kept here because the failure is instructive and the pre-registered check is what caught it. Design 1b replaced it and works on Haiku. See findings.md.

Design 1: the long-form list probe (FAILED)

Have one call emit many controlled decision points inside one long output.

Write exactly 20 numbered lines. Each line has the form:
  <n>. I ate <15 repeated digits> <fruit>.
Choose <fruit> from: mangoes, pineapples, papayas, kiwis.

Why this works:

Watch for:

Design 1b: the preamble probe (WORKS on Haiku)

One probe per response, placed after ≥220 words of story, so there are no sibling choices for the model to balance against. sample_longform.py --mode preamble.

./.venv/bin/python src/sample_longform.py --model claude-haiku-4-5 \
    --mode preamble --digits 9 --calls 4500 --concurrency 12 \
    --out data/$(date +%F)/lf_haiku45.jsonl

Verified on 120 live calls: median 331 output tokens on Haiku, 100% parse, 4/4 words live, max share 0.60–0.65 per digit. Better dynamic range than the original forced choice and above the 200-token threshold.

Two consequences worth noting:

Design 2: the length sweep — the discriminating experiment

The single most informative variable available, and it is nearly free once Design 1 works. Run the same probe at total output lengths spanning the threshold:

Lines per call Approx output tokens Regime
2 ~25 exempt
8 ~100 exempt
16 ~200 boundary
30 ~380 mandated
60 ~750 mandated, high reliability

Then read the result off the shape:

Note the sweep is confounded with per-line position: line 2 of 60 sits in a different place than line 2 of 2. Analyse at matched line indices across lengths, not pooled.

Design 3: a local ground-truth positive control (revised 2026-08-11)

The deepest weakness of every result so far is that the harness has never detected a real watermark — only synthetic ones it generated itself.

The original plan was to buy Gemini calls and hope. Reading the SynthID source produced something better and free: the reference implementation runs locally, GPT-2 on CPU. See synthid.md. That gives ground truth with known keys, watermarked and unwatermarked text from the same model, and no API budget.

It also produced a result that outranks everything else in this file:

Our repeated-digit context is skipped by SynthID by construction. The scheme does not watermark a token whose preceding 4-gram was recently seen, and a run of five identical digit tokens makes the fruit token’s context a repeat of the one immediately before it. The signal is absent at the exact position we measure.

Fix the stimulus before spending anything. Replace the repeated digits with nine fixed nonce phrases whose tokens are distinct. Same context control, no repetition for a guard to skip, and it strictly dominates the digit for KGW too.

Order: run the local control, quantify the digit-vs-nonce gap on ground truth, then collect.

Design 4: restoring entropy on Opus 5

Still unsolved after six attempts. Under the preamble design Opus 5 improved from 95% to 78% on papayas but left mangoes and pineapples never chosen — 2/4 words on a constant rail, which is still degenerate. Framing is no longer the bottleneck; the word list is.

Order of work

Revised 2026-08-11 after the probes and the SynthID read.

  1. List-format probe done, failed — models round-robin the list. The preamble design (1b) replaced it and works on Haiku.
  2. Switch the stimulus from repeated digits to nonce phrases. Free, and without it a SynthID-class mark is invisible at our measurement point.
  3. Local SynthID positive control — free, GPT-2 on CPU. Confirm our pipeline separates watermarked from unwatermarked ground truth, and measure the digit-vs-nonce gap directly instead of arguing it.
  4. Cluster by call not needed — preamble mode is one probe per call.
  5. Length sweep on Haiku 4.5 (~$5) once the stimulus is fixed.
  6. Only then the full high-n collection.

Nothing paid should happen before step 3 clears.

What would count as a positive result

Stated in advance, so it cannot be chosen after seeing the data:

Anything less is a lead, and gets recorded as post-hoc — see the digit-1 anomaly in findings.md for how easily this project generates those.


Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub