Can Turnitin Detect ChatGPT? What Turnitin's Own Data Actually Says
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Quick answer
Yes, in most cases. Turnitin says its AI-writing detector correctly identifies AI-generated content with a document-level false-positive rate under 1%, and its own Chief Product Officer has said the tool is tuned to catch roughly 85% of AI writing, trading some misses for fewer false accusations (BestColleges, "Testing Turnitin's New AI Detector"). That's a genuinely strong number. It is not the same as "Turnitin is always right," and two major universities have published their own reasons for pulling back on it anyway. Here's what Turnitin's own data says, what independent testing found, and what a single score should and shouldn't decide.
What Turnitin actually claims, in Turnitin's own words
Turnitin doesn't hide its numbers — it publishes them. The company's stated design goal is a document-level false-positive rate under 1%, meaning fewer than 1 in 100 fully human-written documents get incorrectly flagged as AI-generated, specifically for documents where 20% or more of the text is AI-written (Turnitin, "Understanding false positives within our AI writing detection capabilities").
That's the headline. The more interesting number sits one level down: Turnitin's own follow-up post states its sentence-level false-positive rate is around 4% — meaning that within a document, individual sentences highlighted as "AI-written" have roughly a 1-in-25 chance of actually being human-written (Turnitin, "Understanding AI writing detection: False positive rates"). Turnitin explains that this happens most often right at the seam where human and AI writing meet in the same document — 54% of these sentence-level false positives sit directly next to genuine AI text, per the company's own analysis.
Chief Product Officer Annie Chechitelli put the trade-off in plain terms to BestColleges: "We would rather miss some AI writing than have a higher false positive rate. So we are estimating that we find about 85% of it. We let probably 15% go by in order to reduce our false positives to less than 1 percent." That's Turnitin, in its own words, confirming the detector is deliberately built to under-flag rather than over-accuse — which answers "can it detect ChatGPT" with a real yes, alongside an equally real "not every time, on purpose."
Turnitin also keeps updating the underlying model rather than shipping it once. Its public release log shows detection updates in both October 2025 and February 2026 described in Turnitin's own words as improving recall "while maintaining a low false positive rate," plus a May 2026 update specifically for Spanish-language detection (Turnitin, AI writing detection model release notes) — evidence the 85%-catch-rate figure from 2023 isn't a permanently fixed number, in either direction.
What happens when institutions test it themselves
Two research universities ran their own numbers in the past year and reached different conclusions than Turnitin's marketing copy alone would suggest.
The University of Waterloo discontinued Turnitin's AI-detection feature in September 2025 after its own Instructional Technologies and Media Services group tested it internally. Waterloo's public rationale cites three independent research papers questioning detector reliability generally, and states plainly that "in more than one instance the product flagged human written text as 100% generated by AI" during its own testing (University of Waterloo, "Discontinuing use of AI detection functionality in Turnitin").
Washington State University went further in February 2026, cancelling its Turnitin AI-detection contract outright. The provost's memo does the arithmetic explicitly: in Fall 2024 alone, Turnitin analyzed 148,547 assessments at WSU — and even accepting Turnitin's own claimed 1% false-positive rate at face value, that implies roughly 1,485 human-written assessments were likely flagged as AI-generated in a single semester. The same memo reports that 33% of academic-integrity hearing board cases involving AI allegations between 2023 and 2025 ended in a finding of "not responsible," specifically because AI-detection output had been submitted as the only evidence (WSU Office of the Provost, "Cancellation of Turnitin AI Detection software").
Neither university is arguing Turnitin never catches AI writing — both memos acknowledge it catches a real share of it. The argument is narrower and more useful: a headline accuracy number describes performance across a large batch of documents, not certainty about any one specific submission, and treating a single score as sufficient evidence for a misconduct finding is where both schools drew the line.
So does it detect ChatGPT specifically?
Independent academic testing backs up that Turnitin can distinguish ChatGPT output from human writing at a meaningfully high rate. A 2025 study published in Acta Neurochirurgica ran 1,000 texts — a mix of pre-ChatGPT human writing and abstracts generated by GPT-3.5, GPT-4, and GPT-4o — through three AI-output detectors and found the detectors "effectively distinguished AI-generated content from human-written texts," with area-under-curve scores between 0.75 and 1.00 across models. The same study's conclusion is the part that matters for a headline claim of "detects ChatGPT": "none of the detectors achieved 100% reliability" (Erol et al., "Can we trust academic AI detective? Accuracy and limitations of AI-output detectors," PMC).
That's the consistent shape of every credible dataset on this question: strong detection of raw, unedited ChatGPT output, a real (if small) false-positive rate on genuine human writing, and a much murkier picture whenever the text has been lightly edited, mixed with human writing, or run through a paraphrasing pass. None of that makes Turnitin (or any detector) useless. It does mean "Turnitin detected 82% AI" is information for a conversation, not a verdict to act on alone — which is exactly the position Turnitin's own chief product officer, Waterloo, and WSU all land on independently.
What to actually do with a score like this
- Treat a mid-range score as a prompt, not proof. Numbers in the 20–50% range sit exactly where every source above says confidence is lowest — that's the range to ask questions about, not act on.
- Ask for process evidence before assuming intent. Draft history, version timestamps, and research notes resolve ambiguous cases faster and more fairly than a percentage does, in either direction.
- Don't rely on one tool's read of one submission. Since detectors disagree with each other in documented testing, a second independent read on the same text is a real check, not overkill.
That last point is exactly what a free check is for. TheChecker.AI's demo gives you a second, sentence-level read on any text — up to 1,000 characters, no signup needed — so you can see where a detection score is confident versus where it's sitting in that murky middle ground before treating any single number as an answer. If you want the underlying mechanics of how any detector, including ours, generates that score in the first place, how AI detectors actually work breaks down perplexity and burstiness in plain terms. And if you're weighing detection tools generally, our own accuracy page publishes the same kind of transparent, tested figures Turnitin does — including where a single score should stop being treated as a verdict, as we've written about before. If you're specifically comparing TheChecker.AI against Turnitin for classroom or team use, our side-by-side breakdown covers that directly.
FAQ
Is Turnitin's AI detector accurate? Turnitin's own published data claims a document-level false-positive rate under 1% and an estimated 85% detection rate for AI-generated content, by design trading some missed detections for fewer false accusations. Independent academic testing generally supports strong — not perfect — accuracy on unedited AI text.
Why did some universities stop using Turnitin's AI detector? Not because it never works. The University of Waterloo and Washington State University both cite the real, non-zero false-positive rate — and the harm a false accusation causes a student — as reason enough to stop treating a detection score as sufficient evidence on its own, even at a false-positive rate Turnitin itself calls low.
Can Turnitin be fooled by lightly edited or paraphrased AI text? Turnitin's own release notes describe ongoing updates specifically targeting "AI bypasser" and paraphrasing tools, which implies the detector's confidence drops on edited text compared to raw AI output — consistent with what independent research finds across detectors generally, not a Turnitin-specific weakness.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how educators and teams can use it responsibly.
Related posts

Can AI Detectors Be Wrong? What the Research Actually Shows
Yes — and the honest answer is more useful than a reassuring one. Here's what peer-reviewed research and real 2025 cases show about false positives, who gets flagged unfairly, and how to use a detection score responsibly instead of as a verdict.
Read more
Does Grammarly Count as AI? The Question Every Student and Teacher Keeps Getting Wrong
Grammarly isn't one thing anymore. Part of it is a 15-year-old grammar checker; part of it is a generative AI writing assistant. Academic-integrity policies are only just catching up to that split, and it's costing real students real grades.
Read more
7 Telltale Signs a Piece of Writing Was Generated by AI (and Why They're Not Enough on Their Own)
Repeated transition words, suspiciously even sentence rhythm, and hedge-everything phrasing are real patterns readers have started to notice in AI writing. Here are seven, and why spotting them by eye still isn't the same as knowing for sure.
Read more