Confirming AI Flags Doesn't Work: 5 Ways Schools Try Anyway (None of Them Prove Anything)
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
When a paper gets flagged, the first instinct is to double-check it: run it through another detector, ask the chatbot itself whether it wrote the piece, compare the work to the student's older essays. A peer-reviewed paper published this year in the Journal of Higher Education Policy and Management tested exactly these "verification" steps and found the same result for all five: confirming an AI flag this way doesn't work, because none of them prove anything. Each one just repeats the same guess in a different costume and calls the repetition proof. The research lays out, method by method, why these five common checks fail.
The problem underneath all five methods
A detector score is a statistical guess, not a measurement. That distinction matters for what happens next. Spam filters and medical tests can be checked against a known reality — the email either was spam or it wasn't, the patient either has the condition or doesn't. AI detectors can't be checked that way in the real world, because once a student submits a paper, nobody outside that student actually knows how it was written. Mark Andrew Bassett and six co-authors from Charles Sturt University, James Cook University, RMIT, Macquarie University, Brisbane Grammar School and SAE University College laid this out in "Heads we win, tails you lose: AI detectors in education", published January 29, 2026: "In real-world conditions, no external evidence can conclusively confirm whether a flagged text was or was not AI-generated. Without a known ground truth, validation efforts rely on subjective interpretation or circular reasoning, rather than objective, independent verification."
That single sentence is the reason every "let's double-check it" strategy below runs into trouble. In the August 5, 2026 Inside Higher Ed report charting the broader move away from detector-as-verdict policies, the same paper is cited directly — this isn't one lab's isolated view, it's shaping how institutions are redesigning the whole process this year.
1. Running it through a second detector
The logic feels sound: if GPTZero and Turnitin both flag the same paper, that's two independent opinions agreeing. Bassett et al. call this out specifically as false confidence. Multiple detectors trained on similar signals — perplexity, burstiness, statistical word patterns — aren't independent measurements any more than two thermometers built from the same faulty design are independent readings. As the paper puts it: "Even if all AI detectors agreed that a text was AI-generated, it would be no more validating than asking a group of phrenologists for a diagnosis — such consensus merely reflects shared flaws, not factual accuracy." Two flawed instruments agreeing with each other is not the same thing as one instrument being right.
2. Hunting for "AI tells" in the writing
Once a score comes back high, it's tempting to go looking for supporting evidence: repeated transitions, oddly even sentence rhythm, overly formal vocabulary. We've catalogued these patterns ourselves — they're real and worth noticing. The problem is using them to confirm a detector's verdict after the fact. The paper's term for this is confirmation bias: once staff suspect a student, they start noticing the features they already expected to find, while overlooking the same features in writing they don't suspect. Formulaic structure, tidy paragraphing, and a habit of starting sentences with "Moreover" show up constantly in genuine human academic writing too, especially from students trained to write that way. Treating a stylistic checklist as independent proof "is not only methodologically unsound but also a clear example of confirmation bias, which can lead to wrongful accusations," the authors write.
3. Asking the chatbot if it wrote the essay
This one sounds almost too obvious to need debunking, but it happens. A teacher pastes a suspect paragraph into ChatGPT and asks: did you write this? The paper is blunt about why that doesn't work: "LLMs cannot recognise their outputs... In some cases, an LLM will confidently — but wrongly — assert that it wrote a passage or that a given text is AI-generated. The model's confidence does not equate to accuracy." A chatbot's yes or no here is a fluent guess, not a lookup against some internal log of everything it has ever generated. Those logs don't exist in a form the model can query about itself.
4. Comparing the paper to the student's older work
This feels like the most human method on the list — surely a sudden jump in polish or a shift in vocabulary means something changed. It might. It also might mean the student got feedback, read something new, worked with a tutor, or simply grew as a writer over a semester. The paper flags this as confirmation bias applied retroactively: once a professor suspects AI use, "changes in clarity, structure, or vocabulary may be misinterpreted as evidence of AI use rather than legitimate progress." Writing style is not a fixed fingerprint. It moves for reasons that have nothing to do with AI, and treating every improvement as suspicious punishes the students who are actually getting better at writing.
5. Hidden "trap" prompts and trick instructions
This is the one that went viral this year — instructors embedding invisible text in an assignment (white-on-white font, a nonsensical instruction) designed to trip up anyone who pastes the prompt straight into a chatbot. Faculty at Brown and Alcorn State both went public with hidden-trap catches this summer, and it made for a satisfying story. The paper takes a harder line on the method itself: it "relies on deception, undermines trust between students and staff, and contradicts the principles of fair assessment," and it's a shelf-life technique — "contingent on current shortcomings in generative AI that can be quickly surmounted by training." A trick that stops working the moment models get slightly better at ignoring off-topic instructions was never evidence in the first place. It was a trap that happened to catch some fish this one time.
The wrinkle even "solid" evidence has now
We've written before about version history as your strongest evidence if you're ever accused, and that advice still holds — a document built up over hours of edits is genuinely harder to fake than a single pasted block of text. But it's worth being honest that "harder to fake" isn't "impossible to fake" anymore. The Bassett paper notes that agentic tools like OpenAI's Operator and Deep Research "can generate, edit, and iteratively refine a document over time, mimicking the human drafting process," and Chrome extensions now exist that literally type AI-generated text into a Google Doc character by character, with pauses and typos built in, to manufacture a clean-looking edit history. Researchers at Georgia Tech and Stanford are already working on tools like DraftMarks, which tries to visualize how AI was used in a draft rather than just whether a history exists. The honest takeaway: version history plus dated outlines and notes is still meaningfully better evidence than nothing, but no single artifact — a score, a style match, a revision log — is proof by itself. That's the same principle running through every method on this list.
What actually helps instead
None of this means a detector score is useless. It means the score should be read as an investigative starting point, not stacked with weak corroboration to manufacture certainty. A few things that hold up better:
- Read the sentence-level breakdown, not just the headline number. Our own detection-score guide covers why a single aggregate percentage hides more than it tells you — a document that's 40% AI by one measure could mean one heavily-generated paragraph or four lightly-edited ones, and those are very different situations.
- Ask for the process, not a defense against the score. The paper's own recommendation is that institutions request drafts and revision history in writing, in advance, as part of the assignment brief — not sprung on a student after a flag, and not as a replacement for the detector, but as its own separate line of evidence.
- Remember the burden of proof sits with the institution. "Students should not be required to prove their innocence," the paper states plainly — an inability to produce a perfect paper trail is not itself evidence of misconduct.
- Check your own writing before anyone else does. Run a draft through TheChecker.AI's free demo to see exactly which sentences are driving a score, instead of waiting to find out secondhand what someone else's detector flagged and why.
FAQ
If detectors can't be independently verified, why use one at all? A score is genuinely useful as a signal that something is worth a closer look — we've covered where that signal is reliable and where it isn't. The research problem isn't the existence of a probability estimate. It's treating attempts to "confirm" that estimate — a second detector, a style comparison, a hidden trap — as if they add independent evidence when they don't.
Does running the same text through a detector twice count as verification? No, for the same reason a single detector's repeat result isn't reassuring on its own: a Wall Street Journal editor got three different results running the same op-eds through the same detector multiple times. Consistency within one tool, or agreement across similar tools, tells you the tools share assumptions — not that the assumption is correct.
Is version history worthless now that it can be faked? No. It's simply not airtight, the same way a detector score isn't airtight. Dated outlines, notes taken alongside a draft, and a revision history that shows non-linear editing (going back to fix an earlier paragraph, not just typing forward) are all harder to fabricate together than any one piece is alone. Combine evidence types rather than leaning on a single one.
What should a school do instead of these five methods? The paper's own recommendation is a shift toward assessment design — oral components, in-class writing, disclosure-tiered assignments — over after-the-fact detection and corroboration. We've covered how several major universities are already making that shift in response to exactly this kind of research.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
Does NIH Actually Detect AI in Your Grant Application?
NIH runs an undisclosed AI detector on grants with real clawback power. NSF just asks for disclosure. A new PNAS study shows what each produces.
Read more
Substack's New AI Detector Is Already Flagging Real Writers. Here's the Actual Lesson.
Substack's new Pangram-powered detector is flagging real human writers. Here's why a strong detector still does that, and what a score actually means.
Read more
LinkedIn's New "Seems Like AI Slop" Button Is a Vote, Not a Detector
LinkedIn's new AI-slop flag button is a crowd vote, not a detector. Here's the accuracy gap real detectors close that a click never can.
Read more