Skip to content
WatermarkAudit

1 October 2026 · 2 min read

Does paraphrasing beat an AI detector?

Usually it does — and what that proves is something about the detector, not about who wrote the text.

Usually, yes. In the paper that tested it properly, running model-written text through a paraphraser dropped one detector's hit rate from 70.3% to 4.6% without changing what the text said.

That number is the interesting part, and it is not really about paraphrasing.

What the research found

Krishna and colleagues built a paraphrasing model called DIPPER and pushed AI-written passages through it, then re-ran the detectors. At a fixed 1% false-positive rate, DetectGPT fell from catching 70.3% of the passages to 4.6%. GPTZero and OpenAI's own classifier degraded too. Later work showed that paraphrasing the output a second and third time defeats even the retrieval-based defences built to resist it.

The reason is simple. A detector like that scores style: how predictable each word is given the ones before it. Paraphrasing changes the words, so the score changes. The origin of the text did not move at all — only the measurement did. A thing you can switch off by rewording is measuring the wording, not the author.

What the detector vendors do about it

Turnitin added AI-paraphrasing detection in July 2024, meant to flag text that looks like it has been through a paraphrasing tool. That is a guess stacked on a guess, and no independent accuracy figures for it have been published, so nobody outside Turnitin can say how well it works.

The statistical watermarks hold up better, because the signal sits in hundreds of token choices rather than in any one phrase. Google's SynthID and the mark Anthropic began applying to Claude's output in August 2026 both work that way. Even they are not immune — the same paper evaded a watermark detector with the same paraphraser — and neither is something you can check: Google's SynthID Detector takes text but is still waitlisted for researchers and newsrooms, and no public detector for Claude's mark exists at all.

So a low AI score means the text no longer looks like what the detector was trained on. It never meant more than that, and a high one never proved authorship either.

What you can check is mechanical: the detector here finds hidden characters and tells you what they are, and AI watermarks explained sets out what the real marks are and who can read them.

Check something yourself

Every tool here runs in your browser and is free. See what your files and text actually carry.

All posts