Back to Blog
Detection Methods & Evidence 7 min read

Can AI Watermarks Hold Up in Court? A New Forensic Study Says No

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

A torn legal document dissolving into scattered ink-wash fragments, its watermark pattern breaking apart across a jagged tear, layered paper-cut diorama style

Quick answer

A July 2026 study built the first framework for testing AI text watermarks against actual courtroom evidence rules, not just lab accuracy. The result: none of the three watermarking methods tested survive a basic paraphrase. Researchers ran 846 valid paraphrase attempts against KGW, Unigram, and Google's SynthID-Text. Every single detectable KGW and Unigram watermark vanished after paraphrasing, a 100% removal rate. SynthID lost its signal in 98.3% of cases. The paraphrased text kept its original meaning, which is exactly why a defense attorney could argue it isn't evidence tampering at all. Watermark detection, on its own, is not the kind of evidence a court can currently rely on.

What the study actually tested

Saifur Rahman Tamim and Amir Labib Khan, computer scientists at Northern University Bangladesh, published "AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation" on arXiv in July 2026, with a version submitted to the AAAI/ACM Conference on AI, Ethics, and Society. Their question was narrow and specific: does watermark detection produce evidence that would survive a Daubert hearing, the standard U.S. federal courts use to decide whether scientific evidence is reliable enough to hear at all?

Daubert asks five things: is the method testable, has it been peer-reviewed, is the error rate known and stable, do controlling standards exist, and is it generally accepted. The researchers paired that legal test with NIST SP 800-86, the federal guide for how digital evidence should be collected and analyzed. Nobody had run watermarking through both filters at once before.

They tested three methods that represent different design approaches: KGW and Unigram, which tag text by nudging a language model toward certain word choices, and SynthID-Text, Google's method, published in Nature and already running in production inside Gemini. All three were run through Qwen2.5-1.5B or Gemma-2-9b-it, then attacked with a same- or cross-model paraphraser, filtered so only genuinely meaning-preserving rewrites counted.

The numbers that matter

Before any attack, detection was already shaky. KGW correctly flagged only 9 of 30 pristine watermarked texts, a 70% false-negative rate. Unigram missed 25 of 30 (83% false negatives). SynthID missed 24 of 30 (80%), with 24 of those 30 landing inside its own built-in "uncertain" zone.

Then came the paraphrase test. Among texts the detector had actually flagged as watermarked in the first place, paraphrasing erased the signal completely for KGW and Unigram: 100 out of 100 detectable KGW texts lost their mark, 41 out of 41 for Unigram. SynthID held on slightly better: 58 of 59 detectable texts lost theirs, a 98.3% removal rate. Multiply the misses together and a court is left with an effective detection rate of roughly 0% to 0.3% after one round of paraphrasing that a defense team could run through any free tool.

SynthID had a second problem: it flagged 10 of 184 paraphrased, entirely human-written control texts as AI-generated, a 5.4% false-positive rate. One of the falsely-flagged samples was a plain typing-test sentence. A statistical detector that mistakes ordinary human writing for machine output on one try in twenty is not a tool a court, or anyone else, should treat as decisive on its own.

None of this is an isolated result. WaterPark, a separate 2025 evaluation published at EMNLP Findings that tested 10 watermarking methods against 12 different attack types, independently found SynthID's true-positive rate dropping to 0.498 under moderate paraphrasing and 0.232 under translation, and found that a single ChatGPT paraphrase pass pushed every tested method below 30% detection. Two independent teams, different methodology, same conclusion: paraphrasing breaks watermark detection fast.

The researchers measured how much meaning survived the attack, because that's the detail that matters in a courtroom. The retained paraphrases kept a median semantic similarity of 0.83 to 0.84 out of 1.0 across all three methods, using sentence-embedding comparison, with every retained sample above the 0.75 cutoff the researchers set as their bar for "clearly still the same text."

That's the crux of the legal argument. A prosecutor can point to a positive watermark hit. A defense attorney can run the same text through an off-the-shelf paraphrasing tool, hand back a version that says the same thing, and the mark disappears. Nobody destroyed evidence. Nobody scrubbed metadata. The words just got rearranged, the way people rewrite sentences constantly. A judge would have a hard time treating that rewrite as tampering, which means the watermark evidence just stopped existing through completely ordinary means.

Scored against the Daubert factors directly, all three methods passed only two of five: they're testable, and they've been peer-reviewed. All three failed on known error rate and on the existence of controlling forensic standards. That's not a minor gap. Factor 3, the known error rate, is usually the factor that decides whether scientific evidence gets in the door at all.

What this means if you're relying on detection, not watermarks

This study is specifically about watermarking, the practice of embedding a hidden statistical signal at the moment text gets generated. It says nothing about statistical detection methods that analyze text after the fact for perplexity, burstiness, and structural patterns typical of machine output, which is a different technique with its own separate error-rate research. But the underlying lesson generalizes past watermarks: any single automated signal, watermark or detector score, is evidence to investigate, not a verdict to hand down. We've made that argument before, about detector scores getting treated as courtroom-ready proof in does a GPTZero flag hold up in court and about why a single AI detector score should never decide a case on its own. This paper is the watermarking side of the same coin, with harder numbers behind it.

It also matters for anyone assuming watermarking is coming for text the way it's already arrived for images. It isn't, not yet, and this study is a big part of why: read the deeper mechanics in AI watermark detection explained and how that gap plays out for a specific model in Claude's new watermark and why you still need a detector.

If you're an educator, an editor, or an HR reviewer trying to figure out whether a specific document was AI-written, the practical takeaway is the same one this study reaches for courts: don't lean on one signal. Run the actual text through a detector built to flag the patterns AI writing leaves behind, read the score as a starting point for a conversation, and go from there.

FAQ

Does this mean AI watermarks don't work at all? They work under narrow, unmodified conditions. The problem is that "unmodified" isn't a realistic assumption. Any basic paraphrase, the kind anyone can run for free in seconds, breaks all three tested methods almost completely.

Is this the same as AI-text detection tools like TheChecker.AI? No. Watermarking embeds a signal at the moment text is generated, and only works on text from a system that chose to add it. Statistical detection analyzes any text after the fact, looking for patterns common in machine-generated writing, whether or not a watermark was ever applied.

Which watermarking method tested best? None passed the researchers' full forensic-readiness bar. Unigram technically cleared the minimum point threshold in their scoring system, but it also had the worst false-negative rate of the three (83%) and 100% conditional removal under paraphrasing, which the researchers flagged as a case where a passing score still described a practically useless method.

Could courts eventually accept watermark evidence? Possibly, if error rates become stable and documented across real-world conditions, which they aren't yet. The researchers frame this as similar to forensic techniques like bite-mark analysis that were used in courtrooms for years before anyone rigorously tested whether the underlying science held up.

Check what you're actually relying on

A watermark that dissolves under one paraphrase, or a detector score you haven't stress-tested, are both single points of failure. If a specific document matters enough to argue about, run it through a detector that flags the actual writing patterns at thechecker.ai/demo, and see how the score holds up against a second look before anyone treats it as proof.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.

Interested in using TheChecker.AI?

Try it free