ccwatermark

Next steps

0a. Watch for the first post-cutoff model (free, and it outranks everything)

Anthropic states that models launched on or after 2026-08-02 are marked at launch. Every model on the API today predates that date — see news.md. So the first post-cutoff release is the first time this project would have a positive control rather than a hypothesis, and it can be paired with a pre-cutoff sibling on identical stimuli to control for ordinary context-sensitivity far better than a cross-vendor comparison could.

./.venv/bin/python src/models_snapshot.py     # free; diffs against data/models/

Run it on a schedule. When it flags a post-cutoff arrival, collect immediately: that model’s launch-day behaviour is as perishable as the pre-rollout baseline was.

Ranked by expected value. Reordered 2026-08-11 after reading the Code of Practice — see regulation.md. Most still need nothing from Anthropic.

0. Collect above 200 tokens (blocks everything else)

The Code exempts text shorter than 200 tokens from watermarking. Every existing collection is 9–40 tokens, so the whole baseline sits in the exempt regime and its nulls say little about the mandated mark.

The design for this is in next-cycle.md. Much of it has now been run: the list probe failed, the preamble probe works on Haiku, the SynthID positive control passed locally, and the length sweep is done. Read findings.md for what survived.

This needs a new sampler, not a parameter change:

It also breaks the diff strategy’s continuity: a >200-token collection cannot be differenced against data/2026-08-10/. The perishability argument now transfers to this new collection — a pre-activation baseline in the regime that actually matters is still collectable today and will not be later. That makes it the most urgent item in this file.

Note this does not retire the forced-choice red-green design, which remains the only method here that needs no unmarked control. It means running it on long free-form generations instead of 9-token ones, which is a real redesign — the statistic wants repeated identical contexts, and long free-form text does not give you those for free.

1. Scheduled re-run and diff (was first; now depends on 0)

Anthropic is still adding marking to already-released models, so marking status changes over time. Re-running an identical protocol and differencing against an earlier collection detects activation without any documentation from them, and the structure of the shift maps directly onto the scheme: which contexts got boosted gives the green lists, and the ETH repo already ships delta_estimation.py and estimate_context_size to recover δ and the context width h from there.

This is the only route where waiting helps us. It is also the only perishable one — a pre-rollout baseline cannot be collected retroactively, and data/2026-08-10/ is currently the only one we have.

Weekly, ~$2 per cycle:

D=$(date +%F); mkdir -p data/$D
./.venv/bin/python src/sample_redgreen.py --model claude-haiku-4-5 \
    --samples 100 --prefixes 10 --digits 9 --out data/$D/rg_full_haiku45.jsonl
./.venv/bin/python src/analyse_redgreen.py data/$D/rg_full_haiku45.jsonl
./.venv/bin/python src/diff_redgreen.py \
    data/2026-08-10/rg_full_haiku45.jsonl data/$D/rg_full_haiku45.jsonl

diff_redgreen.py is the differencing step, built 2026-08-11 and validated but never yet run against a real second collection — there is only one. It tests per-cell, per-digit and joint proportion shifts by exact permutation, prints a green-list direction read-out when it fires, and measures its own power each run. See method.md.

Sensitivity on the 2026-08-10 baseline, 100 vs 100 per cell, 100 reps per δ:

δ detected
0.00 7% — the rule is an OR of two tests, so α is not 0.05
0.15 6%
0.25 47%
0.35 77% — sensitivity floor
0.50 99%

So an activation at conventional strength (δ ≈ 2.0) would be unmissable, and the diff stays blind below δ ≈ 0.15. That floor is close to the single-run red-green test’s δ ≈ 0.25, which is expected — both are limited by the same per-cell n.

1b. Request expert access to the detection solution (cheap, legitimate)

Measure 2.1.1 obliges signatories to provide their detection solution free of charge, without volume limits, to a list that explicitly includes independent researchers and research institutions. Sub-measure 2.1.2 lets them gate free-form-text detection behind “verified expert users” — the same list qualifies.

This ends with the real detector rather than an approximation, and costs an email. Conditional on Anthropic being a signatory, which is unverified. Pursue in parallel with the empirical routes, not instead of them: access may carry terms, and a detector we cannot publish about is worth less than one we built.

2. C2PA on generated files (free, readable today)

The file side uses an open standard with public tooling, unlike the text side. c2patool 0.27.9 is installed locally. It will not hand over the text scheme, but marking pipelines usually ship from one place, so a manifest may carry a scheme identifier, version string, or claim-generator name — and it establishes the signing certificate chain.

The Code settles where to look. Free-form text gets a single marking layer — the watermark — because “free-form text cannot transport metadata”. So the Messages API is out, not merely unlikely. Containerised text (the Code’s term for text inside PDFs, Word documents or HTML) does require signed metadata, so generated .pdf / .docx / .html downloads are the text-adjacent surface worth probing, alongside .svg / .png / .jpg. Worth probing systematically:

c2patool <file>              # full manifest
c2patool <file> --detailed   # assertions and certificate chain

3. Watermark stealing (dual-use — read the boundary first)

Jovanović et al., ICML 2024 — reconstructs an approximate detector from black-box queries alone. This is the route that ends with us holding a detector, not just a yes/no.

This technique is explicitly dual-use: the reconstruction that yields a detector is the same knowledge that enables spoofing and scrubbing. Only the detector-reconstruction half is in scope here, and nothing whose primary utility is evasion gets published — see scope-and-ethics.md. If that half cannot be pursued without shipping an evasion recipe, it does not get pursued.

The key advantage: it works on free-form text, where Claude has abundant entropy (981/1000 distinct completions on Opus 5), rather than forced choice, where it does not. That sidesteps the exact obstacle that made the red-green test untestable on Opus 5 — and, now more importantly, it is the only route here that natively operates in the >200-token free-form regime the Code actually mandates. That promotes it relative to route 1.

Cost scales with corpus size; the ICML work used ~2.2M tokens. Roughly $12 on Haiku 4.5, $60–100 on an Opus-tier model. Worth holding until route 1 fires — if marking activates we will know which contexts to target, and stealing gets far cheaper.

4. Monitoring for the technical document

Article 50 compels publication, so it should arrive. See the watch list in news.md. Low effort, but it is the only route that depends on them.

Smaller open items


Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub