One in Three New Science Papers Now Reads as AI-Written. Here's What That Number Actually Means
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Quick answer
A recent independent review examining 12,750 arXiv papers has revealed that nearly a third of papers published since ChatGPT's launch seem to be written by machines. Only 0.4% of early submissions had that style before large language models became prevalent. This marks a noticeable change in academic writing. An article composition that reads as AI-written is not the same thing as a false or dishonest report, though. Reports such as these indicate writing patterns across large sets of data, not assessments of individual research. The information could help academics, authors, and editors who encounter similar flags in their own work. The best course of action would be to respond with a closer look, rather than jumping to conclusions about any piece of writing.
What the study actually measured
The analysis, published by the team behind the detector tool Unslop, sampled 25 papers per field per month across ten arXiv field groups from January 2023 through July 2026, plus a control set of 1,600 papers from 2021–2022, before generative AI tools were in wide use. That gives 12,750 papers total. Researchers pulled the first-submitted version of each paper specifically so a 2026 revision couldn't leak modern AI phrasing back into an earlier data point, and they scored the full body text rather than just the abstract, because the team found abstracts consistently understate the AI-writing signal — the same paper can score under 20% on its abstract and over 70% on its body.
During the pre-ChatGPT control years the detector’s flag rate remained flat at roughly 0.4%. A rapid climb followed within months of ChatGPT’s release in late 2022. Over the most recent full quarter it surged in two waves to about 32%. The study notes a peak of nearly 39% in early 2026. Each measurement is accompanied by a 95% confidence interval.
The spread across fields is wide. Computer science leads at about 65%. Mathematics sits lowest, near 0.7% — but the researchers are explicit that this is likely a detector-sensitivity artifact, not proof mathematicians use AI less: strip the equations and theorem-proof scaffolding out of a math paper and there's often little ordinary prose left for a language-based detector to score. A math paper written with heavy AI assistance can still land a low score simply because the text doesn't look like the kind of writing the detector was trained on. The researchers call their own number a lower bound for that reason — the tool also misses some AI-written papers by design, since it was tuned to keep the false-positive rate very low, which necessarily means some true positives slip through uncounted.
Because the data did not spike, this proves the approach is more stringent than merely inserting text into a random online checker and taking a percentage at face value. That particular design choice is the element worth holding on to before treating "a third" as a concrete fact instead of a cautious estimate. To avoid a figure that rises simply due to an overly sensitive instrument, the team set their detection threshold specifically to the pre-ChatGPT control years. If the 2021 and 2022 papers had shown a spike as well, the issue would lie with a twitchy detector rather than with an actual shift. An early experiment that processed older, pre-AI papers with a popular free detector flagged them, highlighting the false-positive problem that renders any single AI-detection score suspect unless accompanied by a documented, tested error rate.
This sits inside a bigger integrity problem, not next to it
The rise in AI-sounding prose isn't an isolated curiosity. It's arriving at the same time as a separate, more concrete problem: fabricated citations. A team at Columbia University built an automated system called CITADEL and used it to scan roughly 2.5 million biomedical papers on PubMed Central published between January 2023 and February 2026. Cross-referencing 125.6 million individual references against Google Scholar, CrossRef, and OpenAlex, they found 4,046 references whose claimed titles matched no publication anywhere — spread across 2,810 papers. The rate rose sharply: about 1 in 2,828 papers carried a fabricated reference in 2023; by early 2026 that had climbed to roughly 1 in 277, with the fabrication rate highest in review articles, the papers that synthesize evidence for clinical guidelines downstream.
That distinction matters for how you read the AI-writing number. A paper reading as AI-written might mean an author ran a fluent draft through an editing pass, translated from a second language, or restructured paragraphs for clarity — none of which is misconduct on its own. A fabricated citation is a different kind of failure: a specific, checkable claim that turns out not to exist, and it's the kind of error a "did this read as AI-written" score can't catch, because a hallucinated reference can sit inside prose that reads as entirely human. The two problems compound each other precisely because they require different kinds of scrutiny. One is a style question; the other is a fact-checking question, and confusing them leads either to over-policing normal editing help or under-checking the citations that actually determine whether a claim is real.
Preprint servers have started responding to the volume problem directly. arXiv's computer science section chair, Thomas Dietterich, announced in mid-2026 that authors uploading papers with what the policy calls "incontrovertible evidence of AI slop" — copied prompt text left in the manuscript, uncorrected model commentary, hallucinated references — face a one-year ban from the server, after which they need a paper accepted through peer review before they can upload to arXiv again. Steinn Sigurðsson, arXiv's scientific director, framed the motivation as one of scale: these aren't rare, high-profile incidents anymore, they're happening constantly across a huge submission volume that's overwhelmingly reviewed by volunteers. The policy drew real pushback too — several researchers worried aloud about whether penalties would extend to innocent co-authors, and arXiv had to clarify that only the uploading author is penalized, and that the peer-review requirement lifts once that author has a handful of papers accepted elsewhere.
Why "a flag" and "an accusation" have to stay separate
The Unslop researchers are clear: their tool is not perfect. A pooled false-positive rate of 0.4% comes from a control sample that holds about 200 papers per field. With that rate, eight papers in the 2,000-paper control set trigger the detector. Those flagged papers are distributed over ten fields. Each field's control rate fluctuates, yet the overall pooled rate holds up. The detector cannot test every possible model, prompt style, or editing workflow that authors use. That is why its miss rate on genuine AI-written papers is an estimate, not a guarantee.
Most importantly: the researchers state plainly that a flag is not proof of authorship. Their tool measures whether text reads as machine-written at a calibrated probability; it cannot tell you whether an entire paper was generated end to end, whether an author used AI only to smooth a rough section, or whether the underlying research and reasoning are sound. That's the same limitation every AI detector carries, ours included — a score describes a statistical pattern in the writing, and the value of running that score is knowing where to look harder, not treating the number itself as a conclusion about the person who wrote it.
This is the same reasoning that shows up across academic-integrity research more broadly. Detectors have measurable, documented false-positive rates that vary by writing style, and research specifically comparing detector output across writer populations has found meaningfully uneven accuracy for non-native English speakers — the same structural problem that shows up in student-essay detection studies shows up in research-writing detection too, because the underlying signal detectors key on (predictability, sentence-length variance, word choice) correlates with fluency, not honesty.
What to actually do with a score like this
So what should a reviewer, editor, or researcher actually do with a score like this, instead of treating it as either a shield or a weapon? A few rules, drawn straight from the researchers' own caveats:
- Treat a high score as a prompt to read closer, not a verdict to act on. A flagged paper deserves a careful look at its citations, its methods section, and whether its claims hold up — not an automatic rejection based on the number alone.
- Check citations separately from writing style. The CITADEL findings show fabricated references can hide inside prose that scores low on any AI-writing detector. Verifying that a cited paper actually exists is a different, and arguably more urgent, check than scoring the prose.
- Know that field and language variance are real. A low score in a notation-heavy field, or from a non-native English writer, tells you less than the same score would from a plain-prose native-English piece. Calibrate your read accordingly.
- Ask, don't accuse. Editors who've navigated this well treat a flagged submission as grounds for a direct question to the author about their process, not a public conclusion.
Curious how the scoring actually looks on your text? Run a paragraph, a draft abstract, or an AI-edited section through TheChecker.AI's free demo. It returns a detailed, sentence-level analysis rather than a single aggregate score. That kind of fine-grained feedback is precisely the sort of detail the researchers say readers should demand before relying on any single number.
FAQ
Does a high AI-writing score mean a paper is fraudulent? No. It means the prose statistically resembles machine-generated text. That can result from heavy AI drafting, but also from AI-assisted editing, translation, or paraphrasing of a human's own ideas — none of which is inherently dishonest. Separately, always check whether the paper's citations and data are real; that's a different failure mode entirely.
Why is computer science so much higher than mathematics in this study? The researchers think it's partly a detector-sensitivity issue, not just an adoption difference. Math papers are dense with equations and formal proof structure; once you strip that out, there's little prose left for a language-based detector to score, so heavy AI use in math writing may simply be harder to catch.
How is this different from AI detection in student essays? The mechanics are similar — both score writing style rather than proving authorship — but the stakes and remedies differ. A flagged student essay usually triggers an academic-integrity process with real consequences for one person; a flagged paper on a preprint server is, per the researchers' own framing, a population-level signal about writing trends, not a targeted accusation against a specific author.
Can I check my own writing the same way? A detector of this kind serves to assess your own writing. TheChecker.AI's free demo returns a sentence-level detection score for the document you upload. The score points out portions that appear to be generated by AI prior to any submission. You receive immediate feedback on each sentence. The service focuses exclusively on your personal draft.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how educators and teams can use it responsibly.
Related posts

Why Does My Essay Get Flagged as AI? What to Actually Do About It
A flag is a probability estimate, not a verdict. Here's what actually triggers false positives, a real case showing the cost of getting it wrong, and the exact steps to take before you panic.
Read more
AI Detection Score Meaning: What That Percentage Actually Tells You
A 62% AI score is not 62% of your text. It's a probability, not a fraction. Here's what detector scores actually measure, and how to read one.
Read more
AI Watermark Detection Explained: Why It Barely Works for Text (Yet)
OpenAI and Google's May 2026 provenance push watermarks images, video, and audio at massive scale. Text watermarking exists too, but independent research already broke it with basic paraphrasing.
Read more