ccwatermark

Detecting text watermarks, before the documentation arrives

View the source and data on GitHub · Contributing · Licence: 0BSD

Providers have begun marking AI-generated text to satisfy Article 50 of the EU AI Act. Anthropic has acknowledged that Claude output is marked but has published no technical detail, promising only “forthcoming documentation”. This repository is an independent attempt to work out what is detectable from the outside, using nothing but ordinary API access and statistics — and to have a dated, pre-documentation record when that documentation lands.

It is a research project. Please read scope and ethics first: we build detectors, never removal or forgery tools.

TL;DR

No watermark detected in Claude’s text — which is not evidence there isn’t one. Every model on the API predates the 2026-08-02 cutoff, so they may not carry one yet; and our tests may be blind to one that is there. Best result: a clean null on Haiku 4.5 in the >200-token regime the EU Code mandates.

What a text watermark actually does

A watermark does not add anything to the text — it re-weights which word gets chosen. Here mango falls 0.50 → 0.45 while durian climbs 0.05 → 0.20. No API exposes these probabilities, so the entire method in this repository is an attempt to recover that shift by sampling the same choice thousands of times.

30-second explainer by Andrew Curran, reproduced with credit. It illustrates the general mechanism every logit-biasing scheme shares, including Google DeepMind’s SynthID Text — though the animation is not among DeepMind’s published SynthID materials, so its origin is unconfirmed beyond Andrew.

Do we understand anything?

Yes — though most of what we understand is about measurement, not about Anthropic. That distinction is the honest headline.

Established about the marking landscape

Established about method — the main contribution

The strongest negative result we hold

A forced-choice probe embedded in 341-token free-form output on Claude Haiku 4.5 — inside the regime the Code mandates, with a stimulus that survives repetition guards — returns a clean null: median p = 1.0000, IQR [1.0000, 1.0000], 0 of 100 bootstraps below 0.05.

That arm is also the one where the confound above is suppressed: a token-keyed watermark reads only the preceding few tokens and is unaffected by an intervening story, whereas semantic context-sensitivity is diluted by it — measurably, with cross-cell spread collapsing from 0.249 to 0.044.

What we do not know

Where to start

Document What it covers
Scope and ethics What this project will and will not do
The problem Why third-party detection is hard
Regulation The 200-token threshold and what the Code requires
SynthID The only readable deployed scheme, and what it implies
Findings Every result, including the ones that turned out wrong
Method How to reproduce, and the API rules the literature predates
Research Literature and prior art
Next steps · Next cycle Where to pick this up

Contributions are welcome — see CONTRIBUTING.

Two things would move this most:

  1. A post-cutoff model. The first model with created_at >= 2026-08-02 is marked at launch by Anthropic’s own statement — a positive control rather than a guess — and can be run against a pre-cutoff sibling on identical stimuli. src/models_snapshot.py watches for it and costs nothing to run.
  2. A negative control: the same protocol against a model documented not to watermark text, which would settle the context-sensitivity confound outright.

Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub