Substack's New AI Detector Is Already Flagging Real Writers. Here's the Actual Lesson.
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer: On July 21, 2026, Substack rolled out an in-app AI-detection feature built on Pangram, one of the most accurate text detectors in independent testing, letting any reader scan a post, Note, comment, or reply and see an estimated human-vs-AI percentage. Within weeks, writers who say they wrote every word themselves started getting flagged, including a 20-year professional writer whose researched, self-edited piece came back "15 percent AI." The story isn't that Pangram is a bad detector. It's that even a demonstrably strong detector produces an estimate, not a verdict, and Substack's own launch post admits exactly that. Reading a score as a probability worth investigating, not a fact worth publishing, is the same standard we build into every score we return.
This month's Substack AI-detector story moved fast. Here's what actually happened. Here's where the pushback has a point. And here's what it means for how you should read any detection score, on Substack or anywhere else.
What Substack actually shipped
Substack CEO Chris Best announced the feature in a post titled "Against Claudefishing", coining a term for passing off AI-written text as your own human work. Best's own post puts the threshold at text "longer than 100 words," and The Verge's reporting matches that figure: any post, Note, comment, or reply over 100 words can be scanned via a "Scan for AI text" option, returning an estimate of how much of the text was written by hand versus with AI assistance. TechCrunch's coverage confirms the same rollout and mechanism, though it cites the cutoff in characters rather than words. Writers can pre-scan their own drafts, add a "How I make this" disclosure statement, or turn detection off entirely for a piece, though turning it off still shows readers a visible "AI detection is unavailable" notice rather than nothing.
Best was upfront about the tool's limits in the same post: "Pangram can only detect whether AI was used to make the text, not whether great human care went into creating it, nor whether AI tools were used as a source." That caveat is the entire story, because almost nobody reading a crisp percentage on their screen internalizes it that way.
Why a good detector still produces this exact backlash
Pangram is not a fly-by-night tool. It's one of the stronger performers in independent academic testing we've covered elsewhere on this blog, and its own self-reported false-positive rate has ranged from roughly 1-in-10,000 documents (per TechCrunch's coverage of Pangram's $9M raise) to 1-in-24,000 on its newest model, a figure independent commentators have cited and linked back to Pangram's own technical writeup without disputing it. Even at the more conservative end of that range, a platform the size of Substack scanning millions of posts will generate real false positives by simple volume, and the writers who land on the wrong side of one don't experience it as a rounding error. They experience it as a public accusation.
That's precisely what happened to writer Wes O'Donnell, who published a first-person account of Substack's detector estimating that roughly 15 percent of an article he researched, drafted, and revised himself was AI-written. His argument cuts to the real mechanism at work: "Resemblance isn't provenance." Large language models were trained on huge amounts of clear, well-organized human writing, then tuned by human reviewers to reproduce exactly that kind of structure. A career technical writer who has spent two decades producing clean, organized prose is going to share statistical fingerprints with a model trained to imitate clean, organized prose. The detector isn't reading his mind. It's reading his sentences, and his sentences resemble the thing the model was built to sound like.
SAN's reporting on the launch added two more data points worth sitting with. The 2025 peer-reviewed paper it cites found that AI detectors disproportionately marked content from neurodivergent writers as suspicious, a pattern that echoes the documented prejudice toward non-native-English authors in prior detector testing. Turnitin previously asserted its false-positive rate stayed under one percent; the Washington Post, reporting independently, found the actual rate topped fifty percent. Self-reported performance numbers from any detection vendor, even a strong one, need this kind of outside scrutiny before anyone treats them as settled fact.
The disagreement isn't really about accuracy
Not everyone thinks Substack overreacted. Writer Andrew Wu defends Pangram's track record in a widely discussed response, pointing out that Pangram has outperformed GPTZero, Originality.ai, and DetectGPT in independent audits since 2024, and that people who claim false positives rarely produce their actual drafts or process notes when challenged. Both sides of this argument can be true at once. Pangram can be meaningfully more accurate than its competitors, and a nontrivial number of real, honest writers can still get caught in its error margin. Those two facts aren't in tension. They're exactly what a low-but-nonzero false-positive rate looks like once you scan enough text.
The actual disagreement, once you strip away the vendor-accuracy debate, is about what a number on a screen is allowed to mean. O'Donnell's sharpest line names it directly: Substack's interface renders an estimate "as a crisp number, and a number does not read like an estimate to a human being scrolling past it." Fifteen percent reads like a measurement. It's actually a classifier's best guess about which statistical neighborhood a piece of writing falls into, based on comparisons to synthetic and human text it was trained on. That gap between what a score looks like and what a score actually is has been the subject of our own guide to reading a detection score, and this launch is a live, high-profile example of exactly the failure mode that guide describes.
What this means if you write on Substack, or anywhere a score might follow you
- A detector score, even from a strong detector, is a starting point for a conversation, not a conclusion. Best's own launch post says as much. The mistake is a platform, a reader, or an editor treating an estimate as a fact because it's rendered as a percentage instead of a sentence.
- Keep your process evidence regardless of what any detector says. Notes, research links, and a revision history that shows real editing, not just a single pasted block, are the strongest counter-evidence to any false flag, on Substack or in a classroom, as we've covered before.
- Check your own writing before someone else's detector does. Run a paragraph through our free demo to see a sentence-level breakdown instead of a single headline number for an entire piece, exactly the granularity that separates a useful signal from a scary-looking verdict.
- Expect this exact story to repeat wherever detection gets bolted onto a platform at scale. The mechanics here (a strong tool, a real error margin, millions of documents, a UI that renders probability as fact) generalize to any platform that ships detection next, not just Substack.
FAQ
Is Pangram a bad AI detector because this happened? No. Independent evidence, including academic testing we've cited elsewhere, places Pangram among the stronger performers in its category. A genuinely low false-positive rate still produces real false positives once a detector is scanning at platform scale. Being one of the more accurate tools doesn't exempt it from that math.
Should I turn off AI detection on my own Substack posts? That's a personal call, but be aware that turning it off doesn't hide the issue. Readers still see an "AI detection is unavailable" label instead of a score, which can read as evasive even when the actual reason is unrelated to how the piece was written.
How is this different from a school flagging a student's essay? The stakes are different but the underlying problem is identical: a statistical estimate of AI-text resemblance gets treated by a human as conclusive proof of authorship. Universities have recently begun to reverse policies that leaned on this kind of score to establish misconduct, as we've written about elsewhere. A large publishing platform adopting the same scoring mechanism runs into the identical gap: the tool measures resemblance, but readers and platforms use it to make categorical calls about who wrote what.
What should actually change based on this? Not the existence of detection tools. The interface and the norms around them. A score presented with real error-margin context, alongside a stated process from the writer, tells a reader far more than a bare percentage ever will, on Substack or anywhere else a detector's output gets treated like a fact.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
LinkedIn's New "Seems Like AI Slop" Button Is a Vote, Not a Detector
LinkedIn's new AI-slop flag button is a crowd vote, not a detector. Here's the accuracy gap real detectors close that a click never can.
Read more
New AI Transparency Laws Just Kicked In. They Don't Cover a Single Word of Text.
New AI transparency laws took effect Aug 2, 2026. An audit found most compliance detectors failed on edited files. Text isn't covered at all.
Read more
Why AI-Generated Job Applications Are Slowing Down Hiring (And What Actually Helps)
67% of HR leaders say AI-written applications have slowed hiring. Volume isn't the fix recruiters need — verification is.
Read more