ccwatermark

The problem space

What is being claimed

Anthropic states that supported Claude models weave “an imperceptible watermark directly into the text itself”, and that because the watermark is part of the text, it survives copy-paste. Separately, generated files of supported types (.svg, .png, .jpg) carry signed provenance metadata following the C2PA open standard.

The two mechanisms are completely different and should not be conflated:

  Text watermark File provenance
Mechanism Statistical bias in token selection Signed metadata attached to the container
Standard Undisclosed, proprietary C2PA, open and published
Inspectable today No Yes, with c2patool
Survives editing Partially — degrades No — stripped by re-encoding
Survives copy-paste of text Yes N/A

Anthropic hedges its own claim: a detected mark “provides a signal that content was processed by Claude, but is not fully conclusive”, and absence of a mark does not establish that content is not AI-generated.

Why this is hard

The scheme is undisclosed. Detection of every known text-watermarking family requires knowing the scheme and usually a key. We have neither.

The API exposes no logprobs. Every powerful detection method wants the token distribution. Anthropic has never exposed logprobs, so the only observable is sampled text. All work here must estimate distributions by repeated sampling, which is why sample counts are in the thousands.

Sampling parameters were removed. temperature, top_p and top_k are rejected with a 400 on current models. We cannot raise temperature to expose the distribution; we get the default and nothing else.

Thinking interferes. On models where thinking is on by default, thinking tokens sit between the prompt and the sampled text. For any test that depends on controlling the tokens immediately preceding a choice, this destroys the mechanism. Thinking must be explicitly disabled.

Modern models are near-deterministic on forced choice. This is the single biggest obstacle found so far, and it is new since the detection literature was written. See findings.md.

Why a mark might be found even if attribution is not intended

Anthropic frames the mark as transparency rather than an attributable signature. Those are not mutually exclusive: any scheme detectable by a third party carries enough signal to support attribution claims by whoever holds the key, and the distinction rests on policy rather than on the mathematics. Treating the scheme as potentially attributable is the conservative assumption, and it is the one this repository works under.

What would count as success

  1. Detection — a statistically sound positive on a known scheme family, with a characterised false-positive rate and sensitivity floor.
  2. Characterisation — recovery of the scheme parameters (context width h, green-list fraction γ, bias δ), which is most of a detector specification.
  3. An independent detector — a tool that answers “was this written by Claude” without Anthropic’s key or documentation.

Only (1) has been attempted so far, and it has returned nulls with known power.


Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub