ChatGPT Detector for Teachers: How to Actually Use One Without Getting It Wrong
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Quick answer
A ChatGPT detector for teachers works only when the score is a signal, not a verdict. Pair it with the student's draft history, their usual writing voice, and an actual conversation — and it saves time, catching real cases. Used alone as a gotcha, it produces false accusations. There's already a paper trail of what that costs: unanswered emails, docked grades, and a due-process argument now working its way through legal scholarship. This piece walks through what a workable classroom process looks like, using real 2025-2026 cases, not the marketing version.
Why this is even a question
More than 40% of surveyed 6th- to 12th-grade teachers used an AI detection tool during the last school year, according to a nationally representative poll by the Center for Democracy and Technology reported by NPR in December 2025. That's a lot of classrooms running a probability score through a disciplinary process that was never designed for probabilities. And it is already producing the predictable failure mode.
NPR profiled Ailsa Ostovitz, a junior at Eleanor Roosevelt High School in Maryland, accused of using AI on an assignment about the music she listens to. The screenshot her teacher sent her showed a 30.76% probability score — not a majority, not even close to the tool's own confidence range, and on a topic Ostovitz says she genuinely loves writing about. She messaged her teacher asking them to try a different detector. No response came, and her grade was docked anyway. Her mother met with the teacher in mid-November; only then did the teacher say they'd never seen the message, and only then did they stop believing Ostovitz had used AI.
Prince George's County Public Schools, the district involved, made the process failure explicit in its statement to NPR: the teacher used the tool on their own, the district doesn't pay for it, and "during staff training, we advise educators not to rely on such tools, as multiple sources have documented their potential inaccuracies and inconsistencies." The tool wasn't sanctioned. The score was treated as proof anyway. That gap — sanctioned guidance versus what actually happens in a single classroom under time pressure — is the real problem a "how to use a detector" guide needs to close.
What the researchers who study this actually say
Mike Perkins, who researches academic integrity and AI at British University Vietnam, told NPR flatly: "It's now fairly well established in the academic integrity field that these tools are not fit for purpose" when used as standalone proof. His own testing found that popular detectors — including Turnitin, GPTZero, and Copyleaks — flagged human writing as AI and missed real AI writing, with accuracy dropping further once the AI text had been lightly edited or paraphrased.
That's not a fringe finding. It's why Turnitin — the company selling the tool many of these districts pay for — publishes its own caveat: AI writing-detection scores "should not be used as the sole basis for adverse actions against a student," and per the company's own guidance, scores of 20% or lower are flagged as less reliable to begin with. When the vendor's own documentation says don't treat this as sole evidence, and a teacher does it anyway under deadline pressure, the failure isn't really about detector accuracy. It's about process.
The due-process argument, plainly
There's a legal dimension to this that rarely makes it into detector marketing copy. Writing in the Harvard Undergraduate Law Review, Maurits Acosta argues that disciplining a student based primarily on an opaque detector score risks violating the procedural due-process protections courts have already established for students. Under Goss v. Lopez (1975), the Supreme Court held that public school students facing suspension have a right to "an explanation of the evidence the authorities have and an opportunity to present his version." Under Dixon v. Alabama (1961), public colleges can't expel a student for misconduct without notice and a hearing.
The practical problem Acosta identifies: a percentage score from a detector is much harder for a student to meaningfully contest than, say, a plagiarism-checker match, which shows the exact overlapping text. A detector score is a probability with an undisclosed methodology behind it. As illustration, Acosta ran George W. Bush's 2001 inaugural address — delivered years before any generative AI model existed — through ZeroGPT and got an 83% "AI/GPT Generated" result. That's not a knock on any one tool's engineering. It's a reminder that a bare percentage, without a sentence-level explanation a student can actually respond to, doesn't hold up to the kind of scrutiny a real hearing requires.
What a workable process looks like instead
Not every district is getting this wrong. Broward County Public Schools — one of the largest districts in the country, enrolling more than 230,000 students — signed a three-year, $550,000-plus contract with Turnitin specifically to check student work for AI use. But Sherri Wilson, the district's director of innovative learning, described the actual use case to NPR in language worth repeating: "The Turnitin tool is something that helps us facilitate conversation and feedback, not grading." Wilson said the district is "totally aware" the tool isn't 100% accurate. It's deployed as a prompt to look closer, not as the disciplinary decision itself.
That's the difference between a district that gets sued and a district that doesn't. A workable teacher process looks like this:
- Treat any score as a reason to look, not a verdict to act on. If a tool's own maker says a 20%-or-lower score is unreliable, and this is the same threshold most tools use internally, don't cite a low score as evidence at all.
- Compare against the student's actual writing history, not a generic baseline. A student who normally writes short, punchy sentences producing suddenly uniform, generic paragraphs is a real signal. A student whose writing has always looked like this producing the same writing is not.
- Give the student the chance to respond before any grade or disciplinary action lands — in writing, with a real deadline for your own reply. Ostovitz's case went wrong specifically at this step: she asked, and no one answered until her mother physically showed up.
- Run a second, independent check if a first score is borderline or disputed. Different detectors weight different signals — perplexity, burstiness, model-specific fingerprints — and disagreement between two tools is itself useful information, not noise to average away.
- Never let the score alone stand in for a conversation. The paper trail in every failed case above — Ostovitz, and the broader due-process argument Acosta makes — is a case where the human step got skipped, not a case where the tool was uniquely bad.
Where TheChecker.AI fits into that process
This is exactly the workflow our own free demo is built to support, not replace. Run a piece of writing through it and you don't get a single number to act on alone — you get sentence-level highlighting of exactly which passages read as AI-generated and which don't, plus which model the writing pattern most resembles. That's the difference between a score a student can't argue with and evidence you can actually walk through together. Read more on how the underlying detection works in how AI text detection works, and see the honest limits of any detector — ours included — in can AI detectors be wrong.
FAQ
Should I fail a student based on a detector score alone? No responsible detector vendor recommends this, including Turnitin, whose own guidance says scores "should not be used as the sole basis for adverse actions." Use a score as a prompt to review the work more closely and talk to the student, not as standalone proof.
What score counts as "high confidence"? There's no universal threshold that works across every tool and every kind of writing. Most vendors flag their own low-end scores (Turnitin flags 20% and below) as less reliable — treat any score in an ambiguous middle range as inconclusive rather than rounding up to "caught."
Is it legal to discipline a student using only a detector score? That's an open and actively contested legal question, not settled law — see the due-process argument above. What's clear is that giving students notice and a real opportunity to respond, per Goss v. Lopez, is the safer and more defensible process regardless of how the legal question eventually resolves.
What should I do if a student disputes a flagged score? Ask to see their draft history if the assignment was done in a shared doc, compare the flagged piece against their earlier work, and consider running the text through a second independent detector. If the score doesn't hold up under those checks, don't let it stand as the official basis for a grade.
Try it on a real assignment
Already running student work through a detector? Check a piece of writing sentence-by-sentence with TheChecker.AI's free demo and see the same kind of evidence Broward County uses to start conversations, not end them. Built for educators who want a workflow, not just a percentage — see the full educator toolkit for what that looks like end to end.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how educators and teams can use it responsibly.
Related posts

Is AI Detection Accurate? What the 2026 Research Actually Shows
Two new 2026 studies and the field's 2023 baseline, compared: false positives are down sharply, but catching fully AI-generated text still varies wildly by tool, and hybrid text defeats almost everyone.
Read more
Does 'Humanizing' AI Text Actually Beat Detection? What the Research Shows
'Humanizer' tools promise to make AI writing undetectable. Peer-reviewed research and the detectors' own published data tell a messier story — here's what actually happens when you run humanized text through a real detector.
Read more
Can Turnitin Detect ChatGPT? What Turnitin's Own Data Actually Says
Yes, most of the time — but Turnitin's own published numbers show a real gap between its accuracy claim and what it admits about false positives. Here's what the company's data says, what two universities found when they tested it themselves, and what that means for a single score.
Read more