Peer Reviewers Are Quietly Using AI to Write Reviews. Detectors Are Catching It.
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
August 2026 saw Nature publish a story about a scientist whose reply to a journal's editorial staff was simply, "Wow, you guys are good!" He was caught because the reviewer used a large language model to write the peer-review report. When the journal's AI detector caught the report, the reviewer admitted using the model. The American Association for Cancer Research, or AACR, ran the AI detection check. AACR had built this screening system after authors asked if their review comments looked machine-written. The investigation found that writers of manuscripts and the reviewers themselves were using AI far more than they had revealed. The difference between claimed and detected AI usage was large enough to change AACR's perception of trust in its own review process.
The reviewer who got caught
Across AACR's ten journals, a regular check runs to spot AI-generated content. The check employs a commercial detection tool from Pangram Labs, chosen after several authors asked if reviewer notes were produced by a chatbot or by a human. Daniel Evanko, director of journal operations at AACR, explained to Nature that most authors flagged by the tool refuse to admit using a chatbot. The one author who confessed is an outlier, not a trend. Nature reported on the program on August 25, 2026.
That single anecdote sits on top of real measurement. An earlier AACR study — first reported by Science in 2025 and later presented in full at the International Congress on Peer Review and Scientific Publication — analyzed 7,177 manuscripts submitted to AACR journals over six months. The detector found that 36% of abstracts contained at least some AI-generated text. Authors disclosed AI use in their manuscript, as journals now require, in only 9% of those cases. Four times as many authors used AI as admitted to it.
What a decade of manuscripts shows
The AACR team didn't stop at one study. A follow-up analysis presented at the Peer Review Congress, covering 46,500 abstracts, 46,021 methods sections, and 29,544 reviewer comments submitted to AACR journals between 2021 and 2024, put numbers on how fast this changed. Before ChatGPT's public release in November 2022, detection of AI-generated text in any of the three categories was close to nonexistent: out of 18,467 manuscripts submitted from 2021 through Q3 2022, only 7 abstracts scored as highly likely AI-generated, and zero methods sections or reviewer comments did. False positives on that pre-ChatGPT baseline were vanishingly rare too — 0.22% on abstracts, 0.07% on methods, 0.03% on reviewer comments — which is the number that actually matters here: a detector that barely ever flags text from before generative AI existed is a detector you can trust when it starts flagging text from after.
By 2024, the picture had flipped. Abstracts: 23% flagged. Methods sections: 10%. Reviewer comments: 4.8%. One detail is easy to miss and worth sitting with: detections in reviewer comments dropped by more than half in the quarter right after AACR's journals explicitly banned AI use in peer review, before resuming their upward climb. A policy change measurably changed reviewer behavior, at least for a quarter, which means reviewers already knew what they were doing carried some risk. The study also found a wrinkle in who gets flagged: authors affiliated with institutions in non-English-speaking countries were flagged more than twice as often as those from English-speaking countries, a pattern the researchers link to using AI as a writing-assistance tool for a second language rather than to generate an entire submission from scratch. Detected AI-generated text was also associated with a substantially higher pre-review rejection rate, and reviewer comments with detected AI use tended to be rated lower by editors overseeing the review.
What journals actually allow
Every major publisher has published a policy on this, and reading a few side by side shows this isn't one exotic case being litigated by a single journal: PLOS's Ethical Publishing Practice policy states plainly that "AI tools cannot serve as reviewers or as decision-issuing editors," and that "any use of AI tools in peer review (e.g. for data assessment, translation, or language editing) must be clearly disclosed to authors in the review form." ACM's RESPECT conference policy draws the same line but names the exception explicitly: reviewers "may not use generative AI tools to write their reviews," but since writing assistants are now built into most word processors, "reviewers may take advantage of such tools to polish the writing in their reviews." Elsevier's own guidance for reviewers permits AI use "in a supportive capacity, for example to improve the language and structure of a review report," provided the reviewer never uploads the actual manuscript into a third-party tool and discloses the assistance.
None of the policies ban AI. Every policy follows the same principle that underpins our scoring system. When an AI simply assists a reviewer who has already reached a decision, the impact differs from when the AI creates the decision outright. Detection tools can find the pattern of an AI that writes the judgment itself, without having to prove intent. Numbers from AACR add evidence that self-reported disclosure badly undercounts actual usage. This undercounting shows up in both directions of the review relationship.
Why this is a genuinely good detection use case
Peer-review reports are close to an ideal target for text-pattern detection, for reasons that have nothing to do with academic integrity specifically. They're pure prose, usually written in one sitting, on a subject the reviewer is supposed to know well enough to have opinions about without needing to look anything up. That combination — familiar text, generated quickly, without heavy multi-round editing — is exactly the profile where perplexity and burstiness signals hold up best, the same statistical groundwork we've covered in how AI text detection actually works. It's a different setting than the grant-application or job-application prose we've written about before, but the underlying mechanism a detector is measuring doesn't change with the genre.
It's also, like every other application of this technology, not proof of anything on its own. A high detection score is a population-level pattern match, not a verdict about one specific person — the same distinction AACR's own team makes when they say a flagged abstract needs a human editor's follow-up, not an automatic rejection. The AACR case is actually a good demonstration of the honest-broker version of that rule in practice: Evanko's team didn't move straight to rejecting every flagged submission. With more than 2,500 AACR submissions flagged for AI-generated abstracts in six months, his own description of the volume problem was blunt — "it's too many to put a human in the loop" for every case — so the publisher is starting with automated follow-up questions to authors rather than snap judgments, treating the score as a trigger for a conversation, not a conviction.
What this means if you edit or review for a journal
Practical lessons emerging from the data caution against banning AI in peer review. Most major publishers have policies that allow AI to support the process. Relying solely on self-reported disclosure is insufficient because authors under-disclose AI use roughly four to one. When reviewers know a ban is in place, they still produce AI-assisted reviews, yet at lower rates than before the ban. That decline indicates a conscious decision rather than an accidental drop. Checking text with an AI detector before accepting it as the reviewer's or author's own work costs only a few seconds. The detector supplies an independent signal that was not available earlier. Consequently, the initial action should be to run a detector before deciding whether a piece of writing is final or not, the same first step we recommend before assuming any piece of writing is settled either way.
FAQ
Can journals actually tell if a peer review was written by AI? Not with certainty, and no publisher claims otherwise. What AACR and similar detection setups can do is flag reviewer comments that show the statistical patterns associated with AI-generated text, at a false-positive rate under 1% on pre-ChatGPT text used as a control baseline. That's a strong screening signal, not courtroom-grade proof about one specific review.
Is using AI to write a peer review against the rules? Using an AI tool to write a peer review is permissible only insofar as it supports language and structure, not the core assessment. PLOS, Elsevier and ACM all endorse this limited assistance, requiring transparent disclosure. Each publisher explicitly prohibits AI from producing the review's substantive judgment. They also prohibit uploading the manuscript itself into any AI platform.
What happens if a reviewer's AI use is detected? Policies regarding detected AI use are not uniform across publishers. Most opt for a case-by-case review instead of instant rejection. AACR flags the AI-use pattern and engages authors and reviewers directly for follow-up. A detection score is never treated as an automatic penalty by AACR. The standard applied is that every AI-detection score serves as a screening trigger but not a final verdict.
Check your own text before you decide what it means
If you're editing, reviewing, or submitting work where AI-assisted text needs a second look, run it through TheChecker.AI's detector before you decide what it means.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
Where AI Writing Actually Concentrates (It's Backwards From What People Guess)
A Rutgers researcher tested 100,000+ documents. People guess AI writes the throwaway stuff. The data says otherwise.
Read more
Newspapers Are Now Running AI Detectors on Every Opinion Submission. The Numbers Don't Agree With Each Other.
Two audits ran the same detector on newspaper opinion pages. US papers: 16% AI-touched. Dutch papers: 42%. Here's why the gap is real.
Read more
How to Actually Evaluate an AI Detector's Accuracy Claims
A University of Chicago study and an FTC order both show why one accuracy number can't tell you if an AI detector actually works for you.
Read more