Newspapers Are Now Running AI Detectors on Every Opinion Submission. The Numbers Don't Agree With Each Other.
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
Two audits ran the same AI detector over hundreds of newspaper opinion submissions, three weeks apart. Semafor scanned 310 guest columns from The Wall Street Journal, The Washington Post, and The New York Times. Roughly 16% came back fully or partially AI-written. AI Report, a Dutch tech-news outlet, scanned 252 submissions across five major Dutch dailies. There, 42% came back fully or partially AI-written. Nearly a third of everything published in de Volkskrant alone tripped the flag. Same detector. Same 80%-threshold rule. Wildly different outcomes, depending on which newsroom's door the AI walked through. A billionaire investor's on-the-record admission triggered both audits: he'd used AI to write a Wall Street Journal op-ed, and the paper ran it anyway. The real finding here isn't about one op-ed. It's that "should we check submissions for AI" now has real operational answers, and every newsroom that tried it drew the line somewhere different.
The op-ed that made newsrooms start checking
On August 25, 2026, hedge-fund veteran Stanley Druckenmiller published "Let the Bond Market Speak" in the Journal, criticizing Treasury Secretary Scott Bessent's bond-market interventions. The piece read like a machine wrote it. "This wasn't liquidity management, it was price management," one line said. Economist Claudia Sahm ran it through Pangram on social media and got a 100% AI-generated score. NOTUS reporter Jeff Stein confronted Druckenmiller directly. He didn't deny it. "Of course I used AI," he said. "I write everything using AI now for the same reason I use a calculator when I do math problems." He pushed back on one point only: the idea that the whole piece was machine output. He said he'd rejected many of the AI's own suggestions along the way.
The Journal's opinion editor, Paul Gigot, defended running it anyway. "AI is a fact of modern life," he told NOTUS. "The question for us is whether what we publish from contributors reflects an author's original argument, and if the author has the standing and credibility to make it." That standard rests on trust in the byline. It doesn't scale to a stranger's freelance pitch. Three weeks earlier, the Financial Times drew the opposite line. Readers had questioned a Harvard professor's column. The FT appended a correction, confirmed AI had condensed the draft, and stated flatly that its "editorial code of conduct specifically prohibits the use of AI in the writing process." Two major papers. Two incompatible answers. Both defended in public within the same month.
What Semafor found when it checked the math
Reporter Reed Albergotti didn't wait for another controversy to break. Semafor ran Pangram over 310 guest opinion submissions published across the Journal, the Post, and the Times over one month. Ten scored at least 80% AI-generated, another 40 came back partially AI-written, roughly 16% of submissions touched by AI in some way. A Pangram spokesperson quoted in the piece put the tool's own false-positive rate at 0.01% on this kind of text, a number Nature's own reporting on the current generation of AI detectors independently backs as unusually low for the category. The individual flags told a messier story than the aggregate number, too. A New York Times guest essay by former CISA director Jen Easterly got flagged partially AI-written, one of twelve Times guest pieces over the same 30 days to trip the tool, despite the paper explicitly banning AI in "developing and drafting guest essays." A Washington Post column by Dartmouth provost Santiago Schnell scored 100% AI, but Schnell's own account tells a murkier story: he wrote the argument himself, he says, and used ChatGPT only to tighten phrasing and check grammar before submission, the exact ambiguous middle ground Pangram's own model struggles with most, separating heavily AI-edited human writing from AI-generated writing outright. Schnell also pointed to Stanford research on detector bias against non-native English speakers; he has a PhD in mathematical biology and is a native Spanish speaker, and that specific bias pattern is well documented in the detection literature, not a one-off excuse.
What AI Report found running the same test on Dutch newsrooms
Dutch tech journalist Alexander Klöpping's outlet AI Report called its own project a direct replication of Semafor's method, aimed at five major Dutch dailies: de Volkskrant, NRC, Trouw, Het Parool, and Het Financieele Dagblad. Out of 252 submissions published between July 27 and August 27, 2026, 49 came back flagged fully AI-generated and another 57 partially. NL Times and Villamedia, the Dutch journalism trade press, both reported the same numbers independently. That's 19.4% fully AI-generated and 42% touched by AI in some way, about six times the fully-AI rate Semafor found across the three US papers. At de Volkskrant specifically, nearly a third of opinion submissions scored above the 80% threshold. The paper had explicitly emailed every contributor that AI use "is not permitted in the writing of submitted articles, except when used as a spell-checker." Editor-in-chief Pieter Klok told AI Report he was "surprised" by the scale, and called the paper's own detection software "far from infallible, both in negative and positive results," the honest caveat any screening tool deserves. AI Report ran a control the same way Semafor did: it tested Volkskrant opinion pieces published in 2022 and earlier, before ChatGPT existed, and every single one came back correctly identified as human-written, the part of a screening pass like this that actually earns its keep, catching a real shift in the submission pool rather than proving anything about one contributor.
Each of the five Dutch newsrooms landed on its own rule once it saw the numbers. NRC requires every sentence to come from the author, reasoning that "thinking precedes writing." Trouw allows correction and translation only, and keeps its own running list of AI-tell phrases. Het Parool draws a line between "text creation," which it bans, and "text checking," which it allows. FD's editor said the numbers made him uncomfortable, but he isn't blocking AI-assisted submissions outright, since a human editor still reads and weighs every piece before it runs. Five newsrooms. Five different thresholds. All built after the fact, not before.
Screening a submission pool is a different job than accusing one writer
It's worth being precise about what these two audits show, because it's tempting to read them as proof that detectors finally work well enough to police individual writers. That's not the lesson. A population-level score measuring a batch of submissions is a different kind of claim than pointing a detector at one person's work and calling the question settled, and Volkskrant's own policy draws that line explicitly: detection software gets used "only on suspicion," a prompt to ask a question rather than an automatic verdict, and the newsroom still takes an author's denial at face value by default. That's the right way to use a screening signal. It's also exactly the gap that got a Wall Street Journal editor accused, after a single detector re-run produced three different scores on the same text. If your own newsroom, agency, or freelance desk is weighing whether to add a detection pass to intake, these two audits are a workable model: run it across a batch to gauge scale, keep a pre-AI control sample to sanity-check the tool against your own archive, and treat any single flag as the start of a conversation, never the end of one.
Run your own draft through the check before an editor does
If you write opinion pieces, guest columns, or bylined analysis for outlets that have started screening submissions, run your draft through our detector before you send it. You'll get a sentence-level breakdown instead of one pass-fail number, the same distinction that separated the newsrooms handling this well from the ones that didn't.
FAQ
Which AI detector did both newsrooms use, and does a high score mean the writer generated the whole piece? Both Semafor and AI Report used Pangram, applying the same rule: a piece is labeled fully "AI" once the detector scores it above 80% AI-generated, and "partially AI" below that threshold but above zero. A high score doesn't settle how the piece got made, though. Several flagged pieces in Semafor's audit, including the Washington Post column by Dartmouth's Santiago Schnell, came from writers who say they wrote the argument themselves and used AI only to edit and tighten the prose. A detection score measures statistical patterns in the finished text. It says nothing about the process behind it.
Why did the Dutch newspapers score so much higher, and should every publication start scanning submissions? Neither audit's authors named a single cause, and a gap this large, six times the fully-AI rate, deserves caution against over-explaining it from two data points; what both agree on is that the gap holds consistently across every outlet each one checked, not one paper skewing the average. As for scanning submissions generally: a screening pass is useful for understanding scale and catching patterns worth a real conversation, the way Volkskrant, the Journal, and the Dutch dailies are all doing now, each in its own way. It stops being useful, and starts causing real harm, the moment a single score gets treated as a verdict against one named contributor instead of a prompt to go ask them about it.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
Where AI Writing Actually Concentrates (It's Backwards From What People Guess)
A Rutgers researcher tested 100,000+ documents. People guess AI writes the throwaway stuff. The data says otherwise.
Read more
How to Actually Evaluate an AI Detector's Accuracy Claims
A University of Chicago study and an FTC order both show why one accuracy number can't tell you if an AI detector actually works for you.
Read more
Peer Reviewers Are Quietly Using AI to Write Reviews. Detectors Are Catching It.
A cancer-research publisher's AI detector found reviewers using AI to write review reports far more than they admitted. Here's what the data shows.
Read more