Should a Single AI-Detector Score Ever Decide a Case?
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
No case decided so far has turned on whether an AI detector's score was correct. It has turned on what happened after the flag. A fresh Communications of the ACM feature (published September 4, 2026) lays out why: at the College of Charleston, a student failed a class on a single AI-detection score with, in Chris Callison-Burch's words, "no conversation." The student's father sued the school and won. Researchers who study these tools, including Callison-Burch (University of Pennsylvania) and Debora Weber-Wulff (HTW Berlin), agree on the same fix: a detector score is a starting point for a conversation, never a verdict by itself. That's the same standard we hold TheChecker.AI to.
What actually happened at College of Charleston
The details are sparse by design — the case predates the wave of higher-profile lawsuits and settled quietly — but the shape of it is now a template researchers point to. A student received a failing grade based on one number from an AI-detection tool. Callison-Burch, who studies detector reliability at Penn and served as an expert witness in this exact case, told the ACM's Samuel Greengard the problem wasn't the tool's accuracy. It was the process: "There was no conversation." No chance to explain, no second signal, no human step between the score and the sanction. The student's father sued. He won (Communications of the ACM, September 4, 2026).
That single sentence — "there was no conversation" — is the whole finding of every AI-detection dispute that has reached a court so far, ours included in a previous look at the Yale lawsuit. The fights aren't really about whether the detector was right. They're about whether anyone treated the score as one input instead of the entire case.
The research says the same thing from a different angle
The same ACM piece leans on RAID, the largest published benchmark of AI-text detectors: over 6 million generations, 11 language models, 8 writing domains, and 11 adversarial attacks, built by Liam Dugan and colleagues at Penn (arXiv:2405.07940). We've covered RAID's specific failure modes before — how a repetition penalty or a homoglyph can quietly wreck a detector's accuracy with nobody trying to cheat. The point in this ACM piece is broader: RAID's authors put it plainly in their own paper, current detectors are "easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models." A 99%-accuracy claim on a vendor's homepage describes one narrow test condition. It does not describe the messy, real text a professor, editor, or hiring manager is actually staring at.
Weber-Wulff, who has run some of the most-cited independent multi-detector studies in the field, frames the underlying problem structurally: "AI detectors typically operate as stochastic black boxes... Only their creators know how they were trained, how they were tuned, how they process text, and how often they err." Her conclusion isn't "detectors are useless." It's that a black box producing one number is never enough evidence on its own to end someone's academic career, byline, or job application — which is exactly why every one of these lawsuits keeps landing on process, not accuracy.
The pattern holds in every case that's gone public
Line up the disputes and the pattern is identical:
- College of Charleston: single score, no conversation, family sued and won.
- Adelphi University: Turnitin flagged a paper as entirely AI-generated; the student's family spent six figures on legal fees before winning (Inside Higher Ed).
- Yale: a GPTZero flag started a 13-count federal lawsuit that, as we detailed separately, turned on the professor's broader case and the school's process, not the score itself.
None of these cases are arguing "the detector's math was wrong." They're arguing "you let a number make a decision that needed a person." That's a due-process failure, not a false-positive-rate failure — and it's a failure institutions can fix regardless of which detector they use.
What researchers actually recommend instead of a verdict
Callison-Burch and Nick Diakopoulos (Northwestern's Computational Journalism Lab) lay out a short, specific list, not a vague "be careful." First: never use a single score as the sole basis for a sanction. Callison-Burch put it directly: "An AI detector should not be used as the sole mechanism for determining whether someone has violated a policy and then sanctioning them." Second: disclose that a detector is in use, up front, not after someone's already been accused. Third: have the conversation before acting. In his words, "If you adopt a no-AI policy and a detector flags text, the next step is to have a conversation with the person. Otherwise, you will accuse innocent people of cheating." Fourth: check convergence, not one score. Diakopoulos recommends examining error rates for different detectors before deploying any of them — reliability improves "if you use several [detectors] and they converge." One tool agreeing with itself isn't corroboration.
This is close to how we've argued a score should be read on this blog before — see our breakdown of what a detection score actually means and why a positive flag alone can't prove anything — but the ACM piece adds something we hadn't been able to cite directly until now: two independent researchers, on record, describing the exact institutional failure mode that keeps losing in court.
What this means if you're staring at your own flagged score
If your writing got flagged, the CACM reporting and the case history both point to the same two moves. First, don't panic at the number alone — a flag is evidence to bring into a conversation, not a verdict that's already been decided. Second, ask what process comes next: is there a chance to explain, to show drafts or version history, to have someone actually look at more than one score? If the answer is "no, the score is the whole process," that's the exact gap every one of these lawsuits has exploited.
The honest version of what a detector is good for: giving you and the people you're talking to a real, specific number to start from — not an automatic verdict. Run your own writing through TheChecker.AI's detector before you're the one explaining a score to someone else, and check our accuracy page for what our 93% figure does and doesn't claim to cover.
FAQ
Has a court ever ruled an AI detector's score was factually wrong? Not directly, in the cases reported so far. Courts and settlements have turned on process — whether the accused got a chance to respond, whether the school followed its own policy, whether a single score was treated as sufficient proof — not on independently re-testing the detector's accuracy.
Does this mean AI detectors shouldn't be used at all? No. Callison-Burch, who has testified as an expert in one of these cases, says detectors remain useful for spotting content at scale, like AI-driven misinformation accounts. The finding is narrower: a score should never be the entire basis for a high-stakes decision about one person.
What should a school or employer actually do differently? Disclose that detection tools are in use, treat a flag as the start of a conversation rather than a conclusion, and check more than one signal before making a decision. That's the same standard the researchers quoted in this piece recommend, and it's the standard behind every honest detection workflow, ours included.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
What Congress's Floor Speeches Reveal About How AI Detectors Spot Machine Writing
A 135-million-word study of Congress shows which vocabulary tells give away AI-written speeches, and why some lawmakers use far more of them.
Read more
Can an AI Detector Catch a Fake Legal Citation Before It Gets You Sanctioned?
A law firm representing a major bank just got its brief struck for fake AI citations. Here's what a detector can and can't catch in a legal filing.
Read more
Why 99.9% of Scientists Who Use AI to Write Don't Disclose It
A 5.2-million-paper PNAS study found 70% of journals have AI writing policies, but just 0.1% of papers disclose using AI. Here's the gap, and the pushback.
Read more