ccwatermark

What the regulation actually requires

Source: Code of Practice on Transparency of AI-Generated Content, EU AI Office, 38pp, in docs/reference/. This is the first document in this project that states technical parameters. Anthropic has still published nothing itself.

Two caveats before using any of this. The Code is voluntary — adherence “does not constitute conclusive evidence of compliance” with Article 50. And we have not established that Anthropic is a signatory. That check is outstanding and everything below is conditional on it. Even so, the Code encodes what the AI Office and its working groups consider the state of the art, so its numbers are the best available prior regardless of who signed.

The 200-token threshold — the operative fact

Glossary, Very short text:

Text that is so short that in many cases it cannot be watermarked, even with a basic level of reliability. At the time of publication of the code, state-of-the-art techniques enable watermarking with at least a basic level of reliability of text as short as 200 tokens […] Until then, “very short text” should be understood as text shorter than 200 tokens.

Sub-measure 1.1.2:

Signatories will ensure that AI-generated or manipulated content is marked with an imperceptible watermark, with the exception of very short text. For free-form text longer than 200 tokens, watermarking still needs to be applied, even though it may have lower reliability compared to that of watermarking very long text.

So marking is not required below 200 tokens, and is explicitly expected to be weaker near that boundary.

This invalidates the regime of every test collected so far

Collection Length vs threshold
fixed_opus5.jsonl 40 tokens (capped) 0/1000 reach 200
fixed_haiku45.jsonl 25–40 tokens 0/1000 reach 200
rg_full_haiku45.jsonl ~9 tokens 22× below

Both samplers hard-code MAX_TOKENS = 40. Every null in findings.md was measured in the regime the Code exempts.

How much that costs us depends on where the marking is applied

An earlier reading of this (2026-08-11, corrected same day) said the nulls were therefore uninformative. That was too strong. The threshold is a statement about detection reliability from a single document, not about whether the embedding happens — and the two come apart:

Implementation Can it condition on output length? Are our 9–40 token samples marked?
Model watermarking — bias applied per token during inference No. At token 5 the model cannot know the output will reach 300. Yes, marked but individually undetectable
Post-hoc watermarking — applied after generation Yes No, skipped as very short text
Entropy- or length-gated model watermarking Partially No

Sub-measure 1.1.2 names both strategies and encourages model watermarking for model providers. If Anthropic does model-level marking, the bias is present in short outputs too — and our tests aggregate thousands of samples rather than judging one document, so they have power a single-document detector does not.

So the existing nulls are real evidence against always-on model-level marking above δ ≈ 0.25, and say nothing about post-hoc or length-gated marking. That is narrower than the original claim and broader than “uninformative”. The experiment that separates these is a length sweep — see next-cycle.md.

The red-green design is the worse offender: it forces a single-token choice from a 4-word list. That is ~9 tokens, and it is also the minimum-entropy case, which is the second reason a well-built scheme would leave it alone.

Any future collection intended to test for the mandated watermark must produce free-form output well above 200 tokens. That is a different sampler, not a parameter change, and it invalidates direct comparison against data/2026-08-10/.

Free-form vs containerised text

The Code splits text in two, and only one half can carry metadata.

Term Definition Marking required
Free-form text “raw sequence of characters with no enclosing structure, schema, or container format… text shown on a website or within a chat” Watermark only, single layer
Containerised text “text embedded within a structured file format… PDFs, Word documents, or HTML files” Two layers: signed metadata and watermark

Measure 1.1: “given that free-form text cannot transport metadata, a single-layer of marking as described in Sub-measure 1.1.2 is considered sufficient to comply with the requirements of Article 50(2) AI Act for this specific type of content.”

This settles the C2PA question for the API. Messages API output is free-form text by definition, so no metadata layer is required or expected there. Probing it was never going to yield anything. Containerised output — a generated .pdf, .docx or .html — is where signed metadata should appear, and that is the surface route 2 should target.

A legitimate route to the detector

Measure 2.1.1 requires signatories to make a detection solution available free of charge, and:

All Signatories will always provide free access to their detection solution, without any restriction on the volume of requests, to competent market surveillance authorities and other regulators, law enforcement authorities, media, fact-checkers, trusted flaggers, independent researchers, educational and research institutions, and civil society organisations.

Sub-measure 2.1.2 permits access for free-form text watermarking to be restricted to “verified expert users” precisely because it is less reliable — but the same list of qualifying parties applies.

This project is independent research. Requesting expert access is a legitimate, non-adversarial route that ends with the real detector rather than an approximate one, and it is cheaper than every empirical route in next-steps.md. It should be pursued in parallel, not instead — an access grant may carry terms, and a detector we cannot publish about is worth less than one we built.

What the robustness requirements imply about the scheme

Measure 3.3 requires marking to survive, among others:

Homoglyph and character-insertion robustness are further confirmation that the mark is not carried in the characters — consistent with the zero-carrier result in findings.md. Paraphrase and translation robustness is a hard constraint: a plain token-level green-list scheme degrades badly under paraphrase, so either the requirement is being met only partially (“as far as technically feasible” appears throughout), or the scheme has semantic rather than purely lexical structure. Worth holding as an open hypothesis rather than a conclusion.

Dates and a future opening

2 February 2027 — signatories must implement a detection interoperability solution. One permitted option is a signpost:

a publicly readable signpost or other interoperable mechanism in the AI-generated or manipulated content that will signal to the public which detection solution to use

Glossary: “Openly disclosed marking technique within the content… may be implemented, for instance, as an imperceptible machine-readable mark.”

A signpost is openly disclosed by definition. If Anthropic implements one, a publicly documented, machine-readable marker appears in the content. That is a date to watch, and a plausible end to the guessing.

Measure 2.3 also requires detection results to state whether they rest on metadata, a watermark, or forensic detection — so a detector we are granted access to should tell us which mechanism fired.

Open questions this raises


Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub