An AI Detector Couldn't Have Caught Most of 2026's "AI-Writing Hypocrites"
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
The Atlantic published a piece on October 3, 2026 naming a pattern: people whose job is writing about AI, policing AI use, or warning others away from it keep getting caught using AI themselves. Dartmouth Provost Santiago Schnell is the newest example, but the article lists at least four more from the past year and a half: a Western Sydney University integrity officer whose anti-cheating op-ed ran through an AI tool, a media executive whose book on truth in the AI age contained AI-fabricated quotes, an Ars Technica reporter fired for AI-invented quotes in a story about an AI agent, and a Stanford misinformation expert sanctioned by a federal judge for AI-hallucinated legal citations. Run an AI-writing detector on each case and you get two very different outcomes. On some, a detector would have flagged the prose immediately. On others, the detector would have come back clean, because the problem was never the writing style. It was a fabricated fact sitting inside otherwise ordinary sentences.
Five cases, one headline, two different failure modes
The Atlantic's Will Oremus frames all five as versions of the same irony: people entrusted with judgment about AI got caught leaning on it. But the cases split cleanly into two groups once you ask what a detector actually measures.
Group one: undisclosed drafting assistance. Schnell told The Dartmouth he used AI chatbots as "assistive tools," then watched his own explanation of that process score 100% AI-written on Pangram. We covered that case in detail in our post on his denial email. The Sydney case is structurally identical: Western Sydney University's pro vice-chancellor for academic integrity, Professor Cath Ellis, fed roughly 40,000 words of her own material into a Copilot model to produce a Sydney Morning Herald op-ed urging students to "do the work" themselves rather than cut corners with AI, according to Guardian Australia's reporting. Both cases are about writing style and process, exactly what a detector is built to flag. If either Schnell or Ellis had run their drafts through TheChecker.AI before publishing, the sentence-level breakdown would have shown them precisely what a reader (or a student newspaper) would eventually find on their own.
Group two: fabricated content hidden inside the prose. This is the group a style detector mostly can't touch. Media entrepreneur Steven Rosenbaum published The Future of Truth: How AI Reshapes Reality in spring 2026. The New York Times found more than half a dozen misattributed or outright invented quotes in the sections it reviewed, concocted by the AI tools Rosenbaum used while writing a book specifically about AI's threat to factual reliability. Ars Technica fired senior AI reporter Benj Edwards after retracting a story that contained fabricated quotes he'd generated with an AI tool and presented as real, in an article that was, by grim coincidence, about an AI agent that had itself generated a fabricated attack on an engineer. And Stanford professor Jeff Hancock, an expert on deception and misinformation retained by the Minnesota Attorney General's office to defend a deepfake law, filed sworn expert testimony containing citations to academic papers that don't exist. A federal judge excluded his testimony entirely, writing that "the irony" of a misinformation expert falling "victim to the siren call of relying too heavily on AI" was not lost on her.
None of those three pieces of writing would necessarily read as "AI-generated" to a style-and-pattern detector, because the surrounding prose wasn't the problem. A detector built to spot perplexity and burstiness patterns in sentence construction has no mechanism for checking whether a cited paper actually exists or whether a quote was ever said by the person it's attributed to. That's a citation-verification and fact-checking problem, not a writing-style problem, and conflating the two is exactly how a credentialed deception researcher ends up telling a federal court he didn't read his own expert testimony closely enough to notice the sources were invented.
Why this split matters more than the "hypocrisy" headline
The hypocrisy framing is catchy, and it's not wrong, but it buries the more useful distinction. If you're in any role where your writing carries institutional weight, academic policy, journalism, legal filings, published nonfiction, there are two separate risks, and only one of them is something a detector helps with.
The first risk is process transparency: did you use AI to draft or substantially rewrite something you're presenting as fully your own, in a context where that distinction matters to your audience? A detector score is a genuinely useful early-warning signal here, the same way a detector can tell whether a draft reads as AI-polished versus AI-edited, even if it can't settle every edge case on its own.
The second risk is factual integrity: did the AI tool insert a citation, a quote, a statistic, or a name that isn't real, and did you verify it before it went out under your byline or your signature? No detector product on the market, including ours, checks whether a cited case exists or whether a quote was actually said. We made this same point about legal filings with fabricated AI citations: a detector can tell you a brief's prose looks machine-assisted, which tells you where to look harder, but it cannot verify the citations inside that prose. Hancock's case is the same failure mode in an expert-testimony context instead of a legal brief, and Rosenbaum's and Edwards' cases are the same failure mode in publishing and journalism. Three different fields, one identical gap.
Newsrooms are running into a related version of this. Semafor and a Dutch outlet both audited newspaper opinion pages for AI-written content this year and found wildly different rates depending on which newsroom's submission queue they checked, which tells you detection catches style inconsistently even on the problem it's actually designed for. It catches fabricated facts even less reliably, because that was never what it was built to measure.
What actually would have stopped each case
Schnell and Ellis needed disclosure, not detection: say up front that AI assisted the drafting, and the "hypocrisy" framing mostly dissolves, since using an AI tool to tighten phrasing on your own argument is defensible if you're honest about it. Our post on the Sacramento Bee's AI-rewritten bylines makes the same point from a different newsroom: audiences punish hidden AI use far more than disclosed AI use.
Rosenbaum, Edwards, and Hancock needed something detection can't provide: verification. A citation check against the actual database or court record. A callback to confirm a quote was really said. That's editorial and legal process, the kind of check that existed long before generative AI and that generative AI makes more necessary, not less, because a fabricated citation from a language model reads exactly as confident and plausible as a real one.
The practical takeaway
If you're writing something where your credibility is the product, whether that's a policy memo, a published book, a news story, or sworn testimony, run it through TheChecker.AI's free detector before it goes out, especially if an institution or a court might scrutinize how it was produced later. Use the score as a prompt to disclose your process honestly if AI touched the draft. But don't stop there if the piece contains any citation, quote, or statistic you didn't personally verify. A clean detector score tells you the writing style looks human. It tells you nothing about whether the facts inside it are real, and three of 2026's highest-profile "AI hypocrisy" cases happened precisely in that gap.
FAQ
Would an AI detector have caught Jeff Hancock's fabricated legal citations? No. A federal judge excluded Hancock's sworn testimony because the cited academic papers don't exist, not because the prose read as AI-generated. A style detector has no mechanism to verify whether a cited source is real.
Why do people who write about AI policy keep getting caught using AI themselves? Partly because using an AI tool to draft or edit writing is now common enough that it happens even to people whose job is scrutinizing that exact behavior in others, and partly because some of these cases (Schnell, Ellis) are really about undisclosed process rather than deception about content. The ones involving fabricated quotes or citations (Rosenbaum, Edwards, Hancock) are a separate, more serious failure: content that was never verified before publication.
Does loading your own writing into an AI tool for editing count as "writing with AI"? It's genuinely contested, which is part of why these cases spread. Schnell and Ellis both argue that using AI for phrasing and organization on material they wrote themselves isn't the same as having AI generate the argument. That's a real distinction, but it's also exactly the gray zone where a style detector will often read the final output as heavily AI-influenced regardless of how the underlying ideas originated.
What should someone with AI-policy authority do differently? Disclose AI assistance in your own writing before someone else discovers it, and treat any AI-assisted draft as unverified until every quote, citation, and statistic has been checked against a real source. A detector score can flag the first problem. It cannot flag the second.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
The Professor Who Proved AI Detectors Work Is Giving Up on Them
A professor who proved AI detectors work at 88% accuracy is now giving up on them. Here's why the method, not the score, failed.
Read more
A College Provost's Denial Email Also Scored 100% AI. That's the Real Story.
Dartmouth's provost denied using AI to write his op-ed. His denial email tested as fully AI-written too. Here's why that's not a gotcha.
Read more
Detectors Punish Honest AI Editing More Than They Catch Actual Cheating. Here's the Study.
A new Notre Dame study found honest AI editing gets flagged 64-80% of the time, while a full AI draft run through a humanizer gets caught under 4%.
Read more