Back to Blog
Detection Methods & Evidence 7 min read

89% of Biomedical Papers Now Show Signs of AI Writing. Here's the Number Editors Should Actually Trust.

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Layered paper-cut diorama of an open scientific journal with a microscope, DNA helix, and a large indigo 89% numeral cut from paper, evoking a study on AI-written biomedical research.

Quick answer

89% by December 2025. That's the new number for statistical signs of AI-assisted writing or editing, across 1,194,287 PubMed Central papers. Before ChatGPT launched in November 2022, it sat near zero. The team behind the number: Lena Holzwarth, Rita González-Márquez, and Dmitry Kobak, at the Hertie Institute for AI in Brain Health (University of Tübingen) and Ghent University. Full methodology, tables, and limitations sit in the arXiv preprint, posted August 11, 2026; Nature's own coverage ran August 20, and the paper kept circulating on research feeds well into early September. Why build a whole new method? Because their own earlier approach undercounted. So does every other published method. Their earlier number for 2024 was 13-16%. Run the new method on that same year and it jumps to 52% for full papers. Here's what 89% does NOT mean: 89% of papers dishonest. It's a corpus-wide word-frequency signal. Not a verdict on one document. The researchers say this themselves. What the study does pin down, with real precision: where AI writing shows up most in a paper (Discussion and Introduction, far more than Methods), and that researchers from non-native-English-speaking countries show roughly double the estimated usage of researchers from majority-native-English countries. A language-access finding. Not a cheating one. The paper's limits get the same treatment. It assumes non-LLM writing habits from 2023 through 2025 would have tracked the same 2018-2022 trend, absent ChatGPT, and admits that assumption gets shakier the farther the projection stretches, since writers pick up LLM-style phrasing simply from being exposed to it, a limitation the authors state plainly rather than let the headline figure stand unqualified.

What the study actually measured

1,194,287 papers. All English-language. All pulled from the PubMed Central Open Access Subset, spanning 2017 through 2025. Each one carries a standard Introduction, Methods, Results, Discussion structure. Inside that pile, the team tracked 379 "marker words." Words like these, potential, delves. Non-content filler that large language models reach for far more than human writers do. For every marker word, count the yearly usage. Fit a regression line to 2018-2022, the years before ChatGPT existed. Project that line forward. Wherever the projection and the actual observed rate diverge, that gap gets chalked up to LLM-assisted writing or editing. Methods also showed the biggest split between whole-section and cropped-sample estimates, 54% versus 32%. AI-editing wasn't spread evenly through it. Some Methods paragraphs, usually the narrative bits, read as AI-assisted far more than the raw procedural lines do. A sentence like "we used a 2-tailed t-test with alpha set at 0.05" leaves almost no room for a model's stylistic fingerprint. A Discussion paragraph reflecting on what a finding means leaves plenty. That's why a high score in Discussion means far less than the same score turning up in Results, where a fabricated or hallucinated specific would actually matter. Kobak's 2025 number wasn't wrong. It was incomplete. That earlier paper landed on 13.5% LLM usage in 2024 PubMed abstracts using this same frequency-gap idea. What's new here is proof that every frequency-gap method published so far, that one included, only ever measures a floor. Never a real estimate. The quiet assumption baked into all of them: AI-written text hits marker words 100% of the time. It doesn't come close to that. So the team built a tighter formula that drops the assumption, tested it against simulated data with a known right answer, and landed within 2 points across the whole range. Point that corrected formula at the same 2024 abstracts and the number jumps from 13.5% to 31%. Point it at full paper text for December 2025 and it lands at 89%. Here's the part editors can actually put to work: the numbers aren't flat across a paper. Random 255-word crops, controlling for section length, broke down like this for December 2025. Discussion led at 68% estimated LLM usage. Abstracts followed close behind at 67%. Introduction: 59%. Results: 46%. Methods barely moved, sitting at 32%.

The paper's own honest correction is the real story

Here's the detail worth sitting with. This isn't a new team catching an old team's mistake. Kobak co-authored both papers: the 2025 one (13.5%) and this 2026 one (89% for full text, corrected method). The same researcher is publicly saying his own earlier, widely cited number understated reality by a factor of roughly 2 to 6, and explaining exactly why. That kind of self-correction should raise confidence in the new number, not undercut every AI-usage estimate that comes after it. It's also a reminder: no single detection method, corpus stats or a document-level detector, should get treated as a fixed, final number instead of an estimate under a stated method.

Why this isn't the same story as fabrication

That's the same bar we hold our own accuracy claims to. That honesty matters more here than usual, because this study's numbers run so much higher than the last one. This blog already covered a 12,750-paper arXiv-wide study alongside a Columbia citation-fabrication audit, and the same distinction applies: AI-assisted writing and outright research fraud are different problems that just happen to correlate in the news cycle. A Discussion section that reads AI-polished could be entirely honest work by someone who ran a solid draft through an editing tool to smooth non-native phrasing. Fabricated data or invented citations is a different, much more serious problem, and the marker-word method used here can't tell the two apart. It measures vocabulary drift, not truth. The Holzwarth/Kobak paper says this directly: hallucinated citations and outright fraud are separate risks that heavy LLM reliance can make more likely, not the same thing this study measures.

The country-level split matters most here. Researchers publishing from countries without a majority of native English speakers showed an estimated 72% LLM usage in 2025, close to double the 37% from majority-native-English countries. South Korea topped the list at 85%, China at 82%, Taiwan at 80%; the UK sat lowest, 28%. It's tempting to read that gap as one group cheating more than another. That reading is backwards. Research on detector false positives against non-native English writers already points the opposite way: AI tools are closing a real, long-documented language-access gap in English-language publishing, not exposing higher misconduct. Writing in a second language has always meant an unfair disadvantage getting published, and AI-assisted editing is one of the few direct tools available for narrowing that gap.

What this actually means for review and editorial workflows

Nobody's arguing here for a ban on AI in scientific writing, and the study's authors don't either. The point is narrower and more useful: separate AI use that's common and mostly harmless, a polished Discussion section, smoother phrasing, from AI use that would actually be a problem if nobody checked it, a fabricated Results table, an invented citation, a Methods write-up that doesn't match what happened. "Does this sound AI-written" is the wrong question. That's the same logic already pushing journals to screen peer-review reports for undisclosed AI use, and pushing federal funders like NIH to run detection on grant proposals with real enforcement behind it. Read at the sentence level instead of squeezed into one score, a detector check shows exactly which part of a manuscript carries the signal, and why.

FAQ

Does an 89% AI-writing estimate mean 89% of biomedical papers are fraudulent, and why did the same researcher's own earlier estimate jump from 13.5% to 89%? No, it doesn't mean fraud. The measurement tracks word-frequency drift across a huge corpus, not the honesty of any one paper. Editing an honest draft for phrasing and clarity produces the same statistical signal as more concerning uses. And the jump isn't because 2024 usage secretly exploded, it's because the method changed. The 2025 paper measured a floor that assumes AI text hits certain marker words 100% of the time, an assumption this new paper shows understated real usage by roughly 2 to 6x. The corrected method was validated against simulated data with a known answer.

Which section of a paper is most likely to show AI-writing signals, and should high rates from non-native-English-speaking countries count as a red flag? Discussion sections showed the highest usage (68% in December 2025, length-controlled). Abstracts followed close behind at 67%. Methods showed the lowest, 32%, since it's built from procedural specifics that leave less room for a model's fingerprint. And no, high rates from non-native-English-speaking countries aren't a red flag. The authors frame it as AI tools helping close a real, long-documented publishing disadvantage, not evidence of higher misconduct. Research on AI-detector false positives disproportionately flagging non-native English writing backs the same conclusion from the other direction.

Check one manuscript instead of reasoning from an average

Want to check one specific manuscript instead of reasoning from a corpus-wide average? Run the text through our detector. Read the sentence-level breakdown. One score never tells you which paragraph actually carries the signal.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.

Interested in using TheChecker.AI?

Try it free