Back to Blog
Detection Methods & Evidence 7 min read

Can You Check If a Research Paper Is AI-Written Before You Cite It?

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Torn academic paper unfolding into layered paper-cut panels with indigo ink-wash highlighting AI-written sections, a magnifying lens revealing narrow text windows beneath

Quick answer

Yes, and researchers are already doing it at scale. Readers of arXiv can now turn on an AI-detection overlay through a mirror site called alphaXiv, powered by Pangram, and see which parts of any paper on the preprint server score as AI-written before they cite it. The tool doesn't give you a single yes-or-no answer. It slices the paper into windows of text, scores each window separately, and rolls those window-level calls into one headline percentage, the same segmentation approach NeurIPS used in 2026 to screen 969 conference submissions and desk-reject 178 of them. That windowing choice matters more than most readers realize: NeurIPS found that shrinking the text window from roughly 300 words down to 100 words cut the share of "certainly AI" verdicts from 42.7% down to 12.7%, on the exact same set of papers. Before you treat any single-paper AI score as a reason to trust or distrust a citation, it helps to know how that score actually got built.

Why researchers started checking papers before citing them

The trigger for a lot of this scrutiny is straightforward: AI-generated text has become common enough in scientific writing that citing something without checking it first is now a real risk. A study out of the Hertie Institute for AI in Brain Health and Ghent University, tracking word-frequency shifts in 1,194,287 PubMed Central papers, found that markers of AI-assisted writing climbed to an estimated 89% of December 2025 papers, up from near zero before ChatGPT's 2022 release. We've covered that finding in depth here, including why 89% doesn't mean 89% of papers are dishonest. Separately, the American Association for Cancer Research started running its own peer-review reports through a detector after authors asked if reviewer comments looked machine-written, and found that far more reviewers and authors had used AI than had disclosed it, a gap we walked through in our piece on AI-written peer reviews. Put those two findings together and the practical question changes. It's no longer just "is AI in science writing common." It's "given that it is, how do I check one specific paper, right now, before I build an argument on top of it."

What alphaXiv's overlay actually shows you

AlphaXiv is a community mirror of arXiv, the physics, math, and computer-science preprint server. Its AI-detection feature, powered by Pangram, lets any reader flip on a "viewer highlights" mode and see, section by section, which parts of a given paper the model scores as AI-generated. That's a meaningfully different tool than a single-number verdict at the top of a page. It puts the same kind of segment-level breakdown in a reader's hands that professional editors already use, letting you see whether a suspicious score comes from the paper's dense literature-review paragraphs (often the most machine-polished section, since it's largely restating prior work) or from its results and discussion, where an unverified claim would actually matter to your citation.

Why the NeurIPS screening episode is the clearest real-world lesson on window size

The clearest public demonstration of how much a detector's window size can move a paper's headline number came out of an academic conference, not a journal. NeurIPS's 2026 Position Paper Track required submissions to be substantially human-written and partnered with Pangram to check compliance across 969 papers. At Pangram's default window size, 43% of papers scored 90% AI or higher, and 28.2% scored a flat 100%, a number the organizers themselves called "surprisingly high." Before acting on that number, they ran a control: pre-ChatGPT papers from a comparable venue (ACM FAccT 2022) scored 0% under the identical settings, which meant the tool wasn't just miscalibrated across the board. Then they tested how much the window size itself was driving the result. Shrinking the analysis window from roughly 300 words down to about 100 words dropped the share of papers scoring 90% or higher from 42.7% all the way to 12.7%, on the exact same submissions, with no changes to the papers themselves. Shrinking further, to roughly 50-word windows, traded away real detection power: on ten known fully AI-written test papers, recall at the "certainly AI" threshold fell to zero. NeurIPS settled on the 100-word window as its working compromise and still desk-rejected 178 submissions, with another 123 asked to prove substantial human authorship.

The lesson generalizes past this one conference. A large text window catches more real AI use, but it also makes a single AI-polished paragraph drag an entire multi-page paper up to a 100% score. A small window narrows the blame to the actual offending stretch of text, at the cost of missing some genuinely AI-written passages entirely. Neither setting is "more accurate" in the abstract. They're different tradeoffs, and the number you see on any given paper depends on which one the tool in front of you is using by default.

What this means when you're the one deciding whether to trust a citation

If you're checking a specific paper before citing it, whether through alphaXiv's overlay or another detector, the practical version of the NeurIPS lesson is this: a high score on a full paper tells you where to look, not what to conclude. A 100% score built from three or four wide windows might mean one badly AI-polished section pulled the whole document up, exactly the "jitter" effect Pangram's own co-founder has described in independent reporting on the tool: a small edit near a window boundary can swing a document's score by more than the edit itself would justify. A paper flagged as heavily AI-involved is worth reading closely, particularly its methods and results sections, rather than either dismissing outright or citing without a second look. And a claim that anchors your own argument deserves the same scrutiny you'd give any source: check whether the specific number or finding you're relying on sits inside a flagged section or a clean one, and whether the paper discloses AI assistance anywhere in its acknowledgments or methods.

None of this means detection tools are unreliable for this purpose. Independent testing from the research group Epoch AI found that Pangram and a competing detector, GPTZero, both scored zero false positives across 495 human-written passages, meaning neither tool mistakenly flagged genuine human academic, blog, or fiction writing as AI-generated in that test. The tools are accurate at telling you where a text statistically resembles AI output. What they can't tell you, on their own, is whether that resemblance changes whether you should trust the finding underneath it. That judgment still belongs to the reader.

FAQ

Can I check a single arXiv paper for AI-written text myself, without any special access? Yes. AlphaXiv's AI-detection overlay is available to any reader browsing a paper mirrored on the site, no institutional subscription required. It's built specifically for this use case: seeing which sections of one document read as AI-generated, rather than getting a single aggregate score with no way to see what's driving it.

Does a high AI-detection score on a paper mean its findings are wrong? No. A detection score measures how closely the writing style in a section matches statistical patterns typical of AI-generated text. It says nothing directly about whether the underlying data, methodology, or conclusions are accurate. A heavily AI-polished paper can report entirely sound findings, and a paper with no detected AI involvement can still contain errors. Treat a high score as a reason to read the methods and results more carefully, not as a verdict on the science itself.

Why did the same NeurIPS papers score so differently under different detector settings? Because the detector doesn't score a whole document at once. It breaks the text into windows, scores each window, then combines those scores into one number. A wider window (NeurIPS's original ~300-word default) means one AI-heavy paragraph can pull an entire multi-page paper up toward 100%. A narrower window (the ~100-word setting NeurIPS moved to) isolates the flag to a smaller, more specific stretch of text, at some cost to catching genuinely AI-written passages that a wider window would have caught.

Check a section, not just a headline number, before you rely on it

Whether you're screening a citation for your own paper or checking a draft you edited with AI before you submit it, a single aggregate score hides exactly the detail that actually matters, and that goes for your own writing just as much as someone else's paper. Run a paragraph or a full draft through our free demo and see a sentence-by-sentence breakdown instead of one number standing in for the whole thing.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.

Interested in using TheChecker.AI?

Try it free