How to Detect AI in Student Essays: A Step-by-Step Process
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Quick answer
Detecting AI in a student essay isn't one action, it's a short sequence: run the essay through a detector for a signal, then check that signal against the student's actual writing history, then have a real conversation before any grade changes. Skip the middle two steps and you get exactly the failure pattern this blog has already documented — a score treated as a verdict, no chance to respond, a grade docked on a number nobody explained. This piece is the workflow itself: what to gather, what order to do it in, and where a detector score actually helps versus where it just adds false confidence.
Why a detector score alone isn't enough to act on
Start with what the score is actually measuring. A detection percentage is a probability estimate about statistical patterns — sentence rhythm, word predictability, structural habits — not a measurement of how the essay was produced. That distinction matters because the same score means different things depending on what you fed the detector.
A 2025 Stanford SCALE-catalogued study tested this directly. Researchers ran 50 human-written essays and 28 ChatGPT-written essays through GPTZero, sorted by length: short (40-100 words), medium (100-350 words), and long (350-800 words). GPTZero caught AI-generated text almost perfectly regardless of length — 91-100% AI-likelihood across every length bucket. Human-written essays told a different story. Short human essays averaged a 35.56% AI-likelihood score, and long human essays averaged 14.75% — both well above the near-zero score you'd want from a working detector on genuine human writing. Medium-length essays (100-350 words) scored the most accurately of the three groups. Across the full sample, GPTZero misclassified 8 of 50 human essays as AI-written — a 16% false-positive rate in that test set. The researchers' own conclusion: GPTZero is "effective at detecting purely AI-generated content" but its "reliability in distinguishing human-authored texts is limited," and educators "should exercise caution when relying solely on AI detection tools."
The practical read for a teacher grading essays of wildly different lengths: a short response to a discussion prompt and a five-paragraph essay don't carry the same detector-score reliability, even from the same tool, even on the same day. That's not a reason to ignore a score. It's a reason to treat any single number as a starting point, not a verdict — the same conclusion we've written about before applies here at the length level, not just the writer-population level.
What to gather before you act on a score
Turnitin — the company selling AI-detection tools to more school districts than anyone else — publishes its own process for what to do after a score comes back elevated, and it does not start with the score. It starts with artifacts. Their guidance for educators tells instructors to gather "previous drafts that clearly demonstrate the student's voice or writing style" before drawing conclusions, then check four things: Is the writing style different from other assignments? Is the student's voice different? Does the work reference classroom discussions or personal details AI can't replicate? Are citations present and correct?
That checklist works because it's testing something a detector genuinely can't see: continuity. A detector scores one document in isolation. It has no memory of what this specific student's writing looked like last week. You do, or you can get it — a shared Google Doc's version history, a learning-management-system submission trail, a rough outline turned in earlier in the process. We've written about why version history beats arguing with a detector: it shows the actual writing happening, timestamp by timestamp, which a probability score about finished text structurally cannot do.
Put those two sources side by side — score plus history — and you get a genuinely useful signal instead of a bare number. A student whose writing has always looked uniform and generic producing more uniform, generic writing is not new evidence of anything. A student whose writing has always had rough edges and specific personal references suddenly turning in flawless, generic prose is worth a real conversation.
The stance some institutions have taken instead
Not every school treats detector output as something to act on at all. The University of Pittsburgh's Teaching Center reviewed the available tools directly, tested Turnitin's AI-detection feature after its 2023 launch, and reached a stricter conclusion than most of this workflow assumes: "current AI detection software is not yet reliable enough to be deployed without a substantial risk of false positives." Pitt disabled the tool for its Turnitin suite and, as of this writing, does not endorse or support any AI-detection tool. Worth saying plainly: a real, current institutional position exists that says don't use these tools for disciplinary decisions at all, full stop, not even as one signal among several.
Where that leaves an individual teacher who still wants some signal, without a full institutional detection policy behind them, is exactly the workflow the rest of this piece is built for: score as a prompt to look closer, never as the finding itself, with your own institution's policy as the actual ceiling on what a score is allowed to justify.
A workable process, step by step
- Go ahead and run the essay through a detector, but don't read that number as a percentage of the document itself. A detector score is a probability estimate covering the whole piece — it's not saying "30% of these sentences are AI." We've broken down what the number actually represents elsewhere on this blog; worth a read before the score does more work in your grading than it's built for.
- Check the essay's length and format before weighting the score. Per the SCALE-catalogued study above, short and long submissions produce noisier false-positive rates than medium-length ones on at least one widely-used detector. A borderline score on a very short response deserves less confidence than the same score on a 400-word essay.
- Pull the student's writing history. Prior assignments, drafts, version history if the work was done in a shared document, an earlier outline. Compare voice, sentence variety, specificity of examples, and whether personal or classroom-specific references show up. This is the step Turnitin's own guidance leads with, and it's the step a bare score skips entirely.
- If the score and the history disagree, that disagreement is the actual finding — not proof of anything on its own. A flagged score against a consistent writing history is weaker evidence than a flagged score against a sudden style shift. Neither one alone should decide the outcome.
- Begin by contacting the student with a written notice that includes a firm date by which you will respond. Follow that notice with a request for the student to explain, step by step, how each claim or illustration was derived. When the work is authentic, the learner is able to detail the logical chain that led to the result. When the product is largely AI output, the learner often fails to outline that chain, and the very failure signals the problem more clearly than any percentage match could.
- Run a second, independent check only if the first score is genuinely borderline or disputed — different detectors weight different signals, and agreement between two independent tools is more informative than either one alone.
That sequence is a compressed version of what we've already covered in more depth on running a workable classroom process and why essays get flagged in the first place — this piece exists specifically to answer "detect AI in student essays" as a how-to, not a why, with the length-bias data point neither of those pieces covers.
FAQ
Does a longer essay make detection more reliable? Not automatically, and not in one direction. The SCALE-catalogued GPTZero study found medium-length essays (100-350 words) scored most accurately, while both short and long human essays showed higher false-positive rates in that test. Length matters, but it doesn't move reliability in a straight line — don't assume "longer is safer."
Should I tell students their essays will be run through a detector? Turnitin's own guidance to educators recommends exactly that: explain the tool and how it factors into evaluation before assignments are submitted, so a flagged score isn't the first time a student hears their work might be questioned. Transparency up front reduces the "gotcha" dynamic that makes disputes harder to resolve fairly later.
What if I don't have access to the student's earlier drafts? Should you find no earlier drafts in a student's folder, ask the student to provide them directly. Acceptable examples include outlines, rough drafts, note collections, or any artifacts that map the development of the final piece. When an assignment omits a requirement for interim submissions, the resulting absence points to a procedural shortfall to be fixed for upcoming work. Do not interpret the mere lack of drafts as evidence of AI use on its own.
Is there a detector that doesn't have this length problem? Every detector on the market has documented failure modes; the specific SCALE-catalogued numbers above are for one tool and one small test set, not a claim about detectors generally. That's exactly why a score should never be the only input — run a sample through our free demo to see a sentence-level breakdown rather than one aggregate number, and compare it against a second tool if the stakes are high enough to warrant one.
Run it before you need it
One flagged submission today or a full semester of grading, the same workflow applies either way. Drop a piece of writing into our free demo and watch the sentence-level highlighting do what a bare number never could: give the student something specific enough to actually respond to.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how educators and teams can use it responsibly.
Related posts

How to Prove You Didn't Use AI to Write Something
Version history is the strongest evidence you have. Here's how to document your writing process before anyone ever asks you to.
Read more
AI Detector False Positives: What the 2026 Evidence Actually Shows
The Authors Guild tested pre-2023 writing and found detectors flagging it as AI. Here's what causes false positives and how to respond.
Read more
Why AI Detectors Are Biased Against STEM Writing and Code
A 280,000-sample study found AI detectors fail on code and unfairly flag disciplined STEM writing. Here's what the research actually found.
Read more