Can You Appeal an AI Detection Accusation? What the UK's Ombudsman Just Ruled
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
Yes, in the UK you can escalate an AI-misconduct finding to the Office of the Independent Adjudicator (OIA), the official ombudsman for higher education in England and Wales, once your university's own appeal process is exhausted. In July 2025 the OIA published four case summaries about AI-related academic misconduct. Three of the four complaints were upheld or partly upheld against the university, and three of the four students were international or second-language speakers. The pattern the OIA keeps finding isn't that the detector was necessarily wrong. It's that universities treated a probability score as proof, skipped steps in their own process, and never asked whether the tool performs worse for the exact group of students it was used against.
Four cases, one pattern
A July 2026 HEPI analysis — from the UK higher-education policy think tank — walks through the OIA's four 2025 case summaries. Read each one and a pattern jumps out fast.
Case CS072504 — Partly Justified. An international student was called to a viva after Turnitin flagged their assignment. The misconduct panel told the OIA the student had admitted, during the viva, to using AI to translate and paraphrase their work. The OIA read the full transcript. The student had actually said they used Google to find synonyms because English wasn't their first language — no admission to AI translation appears anywhere in the record. The university also never asked whether Turnitin's detection might be less reliable for non-native English speakers, despite that being directly relevant to an international student's case.
Case CS072502 — Justified. Turnitin flagged an international student. A zero was given and the student was directed to resubmit after a viva. Prior to the hearing, no evidence had been shown to the student. No opportunity was offered to submit written mitigation. The student's claim that Grammarly was acceptable language support was dismissed. No justification was given for the dismissal.
Case CS072501 — Justified. A student with autism was accused after a panel compared their essay's writing style to their other work. Ironically, that same detection software had already flagged one of their earlier essays once, and that essay was later confirmed to be human-written. Once the OIA forced a reconsideration, the university reversed course: no misconduct had occurred.
Case CS072503 — Not Justified. A student was asked to attend a viva because there were concerns about their dissertation. During the viva they told the examiners they had used AI to help write it, and they agreed the AI material was not referenced. The university required students to declare AI use on submission; the student had not. A panel found academic misconduct and gave a mark of zero. The student complained they received no credit for being honest in the viva. The OIA found the complaint Not Justified: the provider considered an appropriate range of evidence (the viva disclosure, draft notes, and written and in-person responses), applied its AI policy, followed its procedures, and the penalty was proportionate.
None of the three upheld rulings say the AI detector was technically inaccurate. What the OIA found, case after case, is that a percentage from a tool got treated as the final word instead of one input into a process that's supposed to include evidence-sharing, a real chance to respond, and a check on whether the tool's known weaknesses applied to this specific student.
Why non-native English writers keep showing up in these cases
Three of the OIA's four 2025 cases involved international or second-language students, and that's not a coincidence produced by a small sample. It matches the accuracy research directly: detection leans heavily on perplexity, a measure of how statistically predictable word choices are. Students writing in a second language tend to use shorter sentences, more common vocabulary, and less idiomatic phrasing — the same surface features that make a detector more confident it's looking at AI output. A Stanford study quantified this: across seven detectors tested on TOEFL essays from non-native speakers, 61% of genuinely human-written essays got misclassified as AI-generated, versus near-perfect accuracy on native-English student writing from the same test.
HEPI's analysis adds the stakes: international students make up 24% of all UK higher-education students and 51% of postgraduates, at a moment when 43% of English universities are projecting budget deficits. Universities depend financially on the exact group of students their detection tools are most likely to misjudge.
This isn't only a UK story
Skepticism about AI detection has led the University of Waterloo, University of Cape Town, and Curtin University to restrict or deactivate the Turnitin feature. Other North American institutions — Northwestern, Georgetown, NYU, and more — have adopted a similar stance. MIT's Sloan School of Management declared the detectors ineffective as sole proof. Instead of defaulting to detection, the school advocated a shift toward redesigning assessments: oral defenses, staggered submissions, and thorough process documentation. These ideas echo the OIA's own principle — verify the learning, not just the output.
The OIA's policy recommendation is blunt: suspend AI detection as primary evidence in misconduct proceedings until the tools are independently validated, and get the sector's regulators to issue joint guidance that a detection score alone can't ground disciplinary action. That's not an anti-detection stance. It's the same point this site makes constantly — a score is a signal worth a closer look, not a verdict that ends the conversation.
What actually gets a case overturned
Reading the three upheld OIA rulings side by side, the same specific failures come up. The evidence wasn't shared before the hearing, so a student can't respond to a percentage they were never shown. There was no chance to submit mitigating evidence in writing — a single viva isn't the same as a documented appeal. Prior drafts, notes, and version history weren't requested or considered, even where the university's own procedures called for it. This is the same gap we cover in how to build a version-history paper trail before you're ever accused — outlines, comment history, and timestamped drafts are exactly the kind of process evidence these rulings say should have been asked for. Nobody checked whether the tool's documented weaknesses applied to this student either. A non-native English speaker, a student with a diagnosed condition affecting writing style — both are populations where the published accuracy research already flags elevated false-positive risk. And the score interpretation itself was wrong: treating a detection percentage as a fact rather than a probability estimate shows up in nearly every one of these cases.
FAQ
Does the OIA only cover AI misconduct cases, or all academic misconduct?
All academic misconduct and other student complaints at its member institutions (most universities in England and Wales). AI-related cases are a growing share, not a separate scheme.
What should a student do if they're accused based on an AI detection score?
Students who face accusations rooted in an AI detection score have the right to request, in writing, a full account of what other evidence was used beyond that score. They may also submit drafts, outlines, or complete revision histories for review. Before approaching any outside body, they must first pursue the university's own internal appeal process. Only once that internal process concludes will the OIA review the complaint.
Does this apply outside the UK?
The OIA is limited to the United Kingdom, yet the pattern it exposes applies worldwide: relying only on a score, without checking for bias, is the exact failure now driving detector policy changes at schools in the US and Canada.
Is Turnitin uniquely unreliable here?
Every OIA case happens to involve Turnitin's detection feature — it's simply the most widely deployed tool in UK higher education, not something the ruling body singled out. The failure the OIA describes (score as sole evidence, no bias check) applies to any single-tool, single-score process, regardless of vendor.
Build the evidence trail before you need it
A score isn't useless. It's simply one input, best weighed alongside drafts, process notes, and a sentence-level breakdown of exactly where a flag landed — not a single headline number. Worried about a flag as a student, or building a fairer process as an educator? Run your own text through TheChecker.AI's free demo and see exactly which sentences triggered a score and why, instead of guessing from a percentage. Same evidence-first standard the OIA keeps demanding from universities.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
What Congress's Floor Speeches Reveal About How AI Detectors Spot Machine Writing
A 135-million-word study of Congress shows which vocabulary tells give away AI-written speeches, and why some lawmakers use far more of them.
Read more
Should a Single AI-Detector Score Ever Decide a Case?
A fresh CACM feature on a College of Charleston lawsuit shows what actually loses in court: not a wrong score, a missing conversation.
Read more
Can an AI Detector Catch a Fake Legal Citation Before It Gets You Sanctioned?
A law firm representing a major bank just got its brief struck for fake AI citations. Here's what a detector can and can't catch in a legal filing.
Read more