What Actually Decides an AI-Cheating Dispute? Five 2026 Cases Say the Same Thing
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
A University of Houston-Downtown senior just won an AI-cheating appeal that took three tries to get right, and his case joins four others that, together, make one thing obvious: the detector's accuracy was never what decided anything.
Mark Pieterson got an F in Music Appreciation last August. His professor said the journal entries and discussion posts he'd submitted were "copy and pasted AI generated." He said no, that's not what happened, and appealed. The same professor who accused him heard the first appeal and rejected it. The department chair heard the second and rejected that too. Only on the third attempt, in front of a student discipline committee of four deans, did anyone actually side with him, citing "concerns regarding the reliability and consistency of the evidence used to support the original finding" (FOX 26 Houston, October 2, 2026). That makes five publicly reported AI-cheating disputes in roughly a year to reach a real decision-making body. In every single one, what decided the outcome wasn't whether the AI detector was right. It was whether the process behind the accusation actually held up.
Five disputes, one dividing line
Here's how that played out differently at Adelphi. A freshman in the university's Bridges Program, which serves students with autism spectrum disorders, found himself accused of leaning on Grammarly's AI features once Turnitin scored his "World Civilization" essay 100% AI-generated. He pushed back, handing over results from two separate AI-detection programs that said the opposite. Adelphi found him responsible regardless. Then the same administrator who'd made that finding turned around and denied his appeal too. So he sued, invoking New York's Article 78 process, a mechanism that lets a court step in and ask whether an institution acted arbitrarily or broke its own rulebook. The answer came back brutal for Adelphi: a judge called the finding "without valid basis and devoid of reason," pointed out the university never gave him an advisor or a real hearing despite its own Code of Conduct demanding both, and ordered the whole thing expunged. Nowhere in that ruling did the detector's math get weighed. What got weighed was the broken appeal. (See Matter of Newby v. Adelphi University, 2026 NY Slip Op 26021, as summarized by the law firm Liebert Cassidy Whitmore.)
Minnesota flipped the script entirely. There, a Ph.D. student got expelled once a faculty committee set his eight-hour qualifying-exam answers beside ChatGPT's own output and spotted what it called "near-identical examples and phrasing." Nothing about the process around that finding broke, though. He had a full hearing in front of a five-member panel. An advocate sat next to him. He got to question witnesses. His eventual appeal reached a vice provost, who let the finding stand. He sued anyway, this time over due process, and this time the university held firm: a federal judge dismissed the claims, finding that detailed notice, real representation, a genuine chance to put forward evidence, and appellate review together covered everything due process requires. (Haishan Yang v. Neprash, D.Minn. Oct. 31, 2025, per the same Liebert Cassidy Whitmore summary.) Same basic accusation Adelphi's student faced. What flipped the outcome was simply whether the process surrounding it held together.
Two more cases haven't reached a verdict, and each is testing a different kind of weak spot. At Yale School of Management, everything traces back to a GPTZero flag on a 2024 exam (our earlier look at this case). That flag alone turned into a 13-count federal lawsuit, now stuck past two years with more than a hundred docket entries and no resolution in sight. The flag itself just opened the inquiry. By the university's own account, what actually drove the suspension was something else entirely, a missing source file the student took months to hand over.
Michigan's case reads nothing like it. An undergraduate there, who lives with generalized anxiety disorder and OCD, says she was accused of AI use three separate times within a single course. Her federal complaint describes the basis for those accusations as "subjective judgments about her writing style and self-confirming AI-comparison outputs" run against nothing more than her own outlines. She's suing under the ADA and the Rehabilitation Act now, and her argument is specific: the university took traits linked directly to her documented disabilities, a formal tone, a rigid structure, visible distress whenever she was confronted, and treated them as evidence of cheating instead of something that needed accommodating. No ruling has come down. But her claim adds something genuinely new to this list. It isn't that the detector called it wrong. It's that the detector's own inputs were shaped by her disability in the first place.
The dividing line, stated plainly
Which brings things back to Pieterson. His is the newest case, and in some ways the simplest. No lawsuit, no lawyer, just a university's own three-step appeal chain doing exactly what a three-step appeal chain is supposed to do, ending with a panel of deans overturning a professor's finding on evidence grounds. It's also a reminder that most of these disputes never reach a courtroom at all, and the ones that stay internal are judged against the same question the ones that go to federal court are: was the evidence behind the accusation reliable and consistent, or wasn't it?
None of the decided cases, not one, turned on whether the underlying AI-detection tool was statistically accurate. We've made this point before looking at a single CACM feature on a College of Charleston case: a detector score is a starting point for a conversation, never a verdict delivered without one. Five disputes later, with real court rulings attached to two of them, that's not a hedge anymore. It's the only variable that's actually decided anything. No court, notably, has gone so far as to say a detector itself got the call wrong, only that the process wrapped around it fell apart or held together.
What this means if you're ever on either side of a flag
If you're a student facing an accusation, the research on false positives already tells you polished, formulaic, or non-native-English writing can trip a detector that's working exactly as designed. What wins a dispute isn't arguing the detector's math. It's documentation: drafts, version history, research notes, and a clear account of your process, the same evidence the Adelphi and Houston-Downtown cases turned on once someone actually looked at it. Our own guide on appealing an AI-detector accusation walks through what a fair process is supposed to look like, drawn from the UK's own ombudsman rulings on exactly this question.
If you're the one running the check, whether that's a professor, an editor, or a hiring manager, the lesson is just as direct. A single number invites exactly the kind of challenge that sank Adelphi's case: no record of what else was considered, no chance for the accused to respond before a decision stuck. A sentence-level breakdown gives you something to actually discuss, not just a score to defend in a room with four deans. Most of these disputes never see a lawsuit at all, Houston-Downtown's included, since they get worked out through a university's own internal appeal chain with no attorney and no court involved. Lawsuits tend to show up once that internal process breaks down on its own, or when a student believes something closer to discrimination, not just a bad call, is driving the outcome, which is the core of what Michigan's case alleges.
See what a flagged sentence actually looks like before a dispute ever starts
A single score reveals almost nothing about which parts of a document actually set it off. Run a passage through our free demo instead, and you'll get a sentence-level breakdown, the kind of granular evidence that separates a defensible flag from a reversible one in every case described here.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
Russian State Media Used AI to Write News Scripts. No AI Detector Would Have Caught the Finished Broadcast.
Anthropic's Sept 2026 report shows AI-polished state media content that detectors can't reliably catch after editing.
Read more
Can You Check If a Research Paper Is AI-Written Before You Cite It?
AlphaXiv now flags AI-written sections in arXiv papers. NeurIPS's own test shows window size alone swings a paper's score from 43% to 13%.
Read more
Why "Polished" AI Writing Gets Flagged More Than Raw AI Text
A Carnegie Mellon study found GPTZero and Pangram score the same model 60-80pp differently based only on instruction-tuning, not authorship.
Read more