ccwatermark

Scope and ethics

This project exists to detect and characterise text provenance marking, using only black-box observation of model outputs we generated ourselves. It is transparency research: Article 50 of the EU AI Act requires AI-generated content to be machine-readable and detectable, and this repository is an attempt to check independently whether that is happening, in advance of the technical documentation that has been promised but not published.

What this project will not do

We do not build, publish, or assist watermark removal, stripping, evasion, or laundering. No tool here has that purpose, and none will be accepted.

We do not build forgery tools either. Making human-written text appear AI-marked is as damaging to the provenance ecosystem as making AI text appear human. Both directions are out of scope.

We do not publish keys or key-recovery material should any ever be obtained.

The EU Code of Practice asks signatories not to make available tools “whose purpose is to circumvent the machine-readable markings added to the AI-generated or manipulated content for transparency” (Measure 1.2). We are not a signatory and the Code does not bind us, but that boundary is the right one and we hold to it.

On dual-use techniques

Some published research in this area — watermark stealing, in particular — is explicitly dual-use: the same reconstruction that yields a detector also yields the knowledge needed to evade one. Where we cite or build on such work, the objective is detector reconstruction for verification. Anything whose primary utility is scrubbing or spoofing stays unpublished, and we would rather drop a line of investigation than ship that.

If a line of work cannot be pursued without producing an evasion recipe, we do not pursue it.

How we interact with the systems studied

Everything here could be reproduced by any customer reading their own outputs.

Responsible disclosure

If this work ever recovers something a provider would consider sensitive — a working detector, a key, a specific vulnerability in a marking scheme — the intended path is to contact the provider first and agree a disclosure timeline before publishing. Findings that are merely negative or methodological, which is everything here so far, carry no such constraint.

On the data published here

All committed data is model output generated by this project from synthetic prompts about invented topics. It contains no personal data, no third-party content, and no user data. Prompts and parsing code are included so that any claim can be re-derived rather than trusted.

Framing the results honestly

A negative result here means “these specific tests, at this measured sensitivity, found nothing”. It never means a provider’s output is unmarked. Several nulls in this repository turned out to be artefacts of the test rather than facts about the model, and they are documented as such. Please read findings with that in mind, and do not cite this work as evidence that any model is unwatermarked.


Part of ccwatermark — independent research on AI text provenance marking. Overview · Scope and ethics · Source on GitHub