How Professors Are Redesigning Assignments Instead of Relying on AI Detectors
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
Professors are not waiting for a perfect AI detector. A September 2026 Nature feature and a parallel report from eCampusNews both show the same shift: faculty are redesigning assignments so a detection score is one signal among several, not the whole verdict. Oral defenses, revision-history requirements, in-class writing, and AI-transparent tasks are replacing "submit an essay and hope the detector catches it." Detection still has a job in that process. It just isn't the whole job anymore.
What changed: two reports, one pattern
Nature's careers team interviewed a dozen educators across computer science, history, sociology, chemistry, and linguistics for a September 1, 2026 feature on how they're redesigning coursework for the AI era. Nearly every one of them had already tried to out-detect the problem and moved past it.
The pressure behind the shift is real. A 2026 Higher Education Policy Institute survey of 1,054 UK undergraduates found roughly 94% use generative AI to help with assessed work, and 12% said they inserted AI-generated text directly into coursework. Separately, a Science study led by Igor Chirikov surveyed more than 95,000 students across 20 US universities and estimated 9% used AI on assignments despite knowing it broke the rules. Neither number is a detector reading — both are self-reported survey data, which is exactly why faculty stopped treating a single detection score as sufficient proof either way.
Two weeks later, eCampusNews ran a companion argument from Honorlock's Jordan Adair: move from "after-the-fact detection" to "verification-first" assessment, building trust checks into the process instead of scrutinizing the final PDF. Adair's piece leans on the 2026 EDUCAUSE Horizon Report, which documents growing institutional distrust in detection tools over accuracy, bias, and procedural fairness — the same distrust we covered when Yale, Vanderbilt, and NYU restricted or disabled AI-detector scores as sole evidence earlier this year.
The research behind the caution
The distrust has a real number attached to it. University of Florida researchers Patrick Traynor, Seth Layton, Bernardo Madeiros, and Kevin Butler tested five leading commercial AI-text detectors for a May 2026 paper at the IEEE Symposium on Security and Privacy. Running detectors against roughly 6,000 pre-ChatGPT security-conference papers and AI-generated clones of the same papers, they measured false-positive rates ranging from 0.05% to 68.6% and false-negative rates from 0.3% to 99.6%, depending on the tool and the text. A simple vocabulary tweak to the AI-generated version was enough to collapse several detectors' accuracy.
"These are not reliable or robust tools to use to measure the problem," Traynor told UF News. "We really can't use them to adjudicate these decisions. People's careers are on the line here." That's the exact framing we've made on this blog before: a detection score is a data point, not a verdict. Traynor's team isn't arguing detectors are useless — they're arguing a single number, used alone, isn't strong enough for high-stakes decisions. Grand Canyon University reached a similar conclusion building its own "verification-centered assessment framework," leaning on faculty review instead of a detector's word.
What professors are actually doing about it
Here's Nicholas Mattei's trick at Tulane: don't fight the AI, use it as bait. Three different chatbots summarize the same sources. Before anyone writes their own sentence, they have to catch which citations got invented along the way. By the time the essay starts, the detecting already happened, and no detector was involved.
MIT's Nikita Bezrukov wanted a different kind of proof. He has students draft everything in Google Docs, because a finished essay only shows him the destination. The revision history shows the route, the same evidence trail covered in how to prove you didn't use AI. That's harder to fake than the destination ever was, and administrators checking the record after the fact end up wanting the exact same thing students are building toward.
A few instructors skipped the page altogether. At San José State, Grazioli grades code 70% on whether it runs and 30% on a one-on-one, out-loud walkthrough of a specific line. UVA's Sessions brought the notebook-only blue-book exam back, three times a term. Princeton's Goldring set a hard ceiling: nobody gets above a C if they can't defend their own argument once someone asks a follow-up question.
Some instructors design tasks a chatbot simply answers badly. Ruben Verborgh at Ghent University keeps his web-development exam open-book and open-internet, AI included, but asks diagnostic questions like "why is this site slow?" that a generic chatbot fumbles without real methodology. Etienne Roesch at Reading assigns analysis of a specific data screenshot; a pasted-in AI answer visibly doesn't match what's actually on screen.
And some professors just downgrade detection instead of dropping it. Jacqueline Evans at Florida International still runs her university's AI-detection tool on submissions, but treats a flag as "some level of corroboration" alongside hallucinated references and a mismatch between a student's draft and final essay. One input among several — the same role Turnitin's own chief product officer has described a detection score as playing.
None of this makes detection obsolete. It changes what a flag is allowed to do on its own.
Where a detector still earns its place
Redesigning every assignment in a syllabus is expensive, and most instructors don't have Grazioli's ratio of TAs or Sessions's small seminar size. For everything short of a full redesign, an instant detection check is still the fastest way to know whether a submission deserves a closer look at all — which is the actual point of Evans's "corroboration" model. A sentence-level read that shows exactly which parts of a document carry an AI-generated signature is a stronger starting point for that closer look than a single whole-document score, and it's the same honest-broker standard we hold our own scores to: a flag opens a conversation, it doesn't end one.
FAQ
Does this mean AI detectors are being phased out? No. Both reports describe detectors being repositioned, not removed — one signal among drafts, revision history, in-class work, and direct conversation, instead of the only signal.
What should I do if my own writing gets flagged under one of these new systems? The same as before: keep your drafts and revision history. See our guide to appealing an AI-detector accusation for the process side.
Are these redesigned assignments realistic for large lecture courses? Partially. Oral exams and Grazioli-style one-to-one checks scale poorly without more graders. That's exactly the gap Marc Watkins at the University of Mississippi's AI Institute for Teachers flagged in earlier reporting: redesigning assessment at scale is labor-intensive, and it shouldn't fall on faculty alone. Lower-effort moves — Google Docs revision history, hallucination-check assignments, pass/fail grading on take-home work — are the ones spreading fastest across departments with the least support to redesign everything at once.
See what a flagged sentence actually looks like before you submit it
A verification-first classroom still needs a fast, sentence-level read of a document, not a single number. Run your draft through our free demo and see exactly what a detector is responding to before a professor — or an appeals committee — ever sees it.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
Harvard's "Two-Way Suspicion": What AI Detection Is Doing to Classroom Trust
A new Harvard report names "two-way suspicion": teachers' AI-dar hunches and student fear, eroding trust without proof.
Read more
Dartmouth Just Authorized an AI Detector for Grading. Its Own Rules Show Why That's Harder Than It Sounds
Dartmouth authorized Pangram for grading with 7 rules: de-identify work first, no grade penalty on a score alone, FERPA-driven.
Read more
What Actually Decides an AI-Cheating Dispute? Five 2026 Cases Say the Same Thing
Five AI-cheating disputes reached real verdicts in 2026. None hinged on detector accuracy — all came down to whether the appeal process held up.
Read more