Back to Blog
Detection Methods & Evidence 6 min read

The Base-Rate Problem: Why a 99% Accurate AI Detector Can Still Be Wrong Most of the Time

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Paper-cut diorama of an old-fashioned balance scale, one side piled high with torn-paper documents marked false accusations, the other side holding just a few correctly flagged documents, indigo and amber ink-wash bleeds

Quick answer

A detector can be 99% accurate and still be wrong more often than it's right when it flags someone. That's not a contradiction. It's basic screening-test math, and a peer-reviewed paper published in Elsevier's Next Research this year finally ran the numbers for AI-text detection specifically. The missing variable isn't the detector's sensitivity or specificity. It's how common undisclosed AI use actually is in whatever population is being screened. Get that number wrong, and a "99% accurate" tool can turn into a coin flip on any single accusation.

The paper: same math doctors have used for decades

In May 2026, economist Panagiotis Tsigaris (Thompson Rivers University) and researcher Jaime A. Teixeira da Silva published "AI detecting AI in academic writing: Why most AI detector findings are false" in Next Research, a peer-reviewed Elsevier journal (DOI: 10.1016/j.nexres.2026.101396). Their framework isn't new. It's the same epidemiological logic used to evaluate any screening test: sensitivity (catches true cases), specificity (correctly clears true negatives), and prevalence, or how common the thing you're testing for actually is in the population you're testing.

The paper's framework is simple: treat an AI detector the way you'd treat a screening test, and a piece of text the way you'd treat a patient. Plug in real-world sensitivity and specificity numbers from existing detector studies, and the conclusion doesn't soften anything: "AI detectors are prone to generate more false accusations than correct identifications," once you factor in how rare deliberate, undisclosed AI misuse actually is across a given pool of essays or manuscripts. None of this is a new framework, either. It's the same epidemiological logic already used to evaluate any screening test, borrowed straight from Ioannidis's famous 2005 argument: a finding's odds of being true depend heavily on how many true findings already exist in a given field, not just on any single test's raw accuracy.

Why "99% accurate" doesn't mean what it sounds like

Here's the mechanism, worked through with round numbers, the same way any epidemiologist would check a screening test before trusting it. Published evaluations consistently show detector sensitivity and specificity around ninety-five percent: the tool correctly flags AI-generated text about 95% of the time and correctly clears human-written text about 95% of the time too. Only about five percent of the items in a typical batch are expected to contain undisclosed AI content, roughly what you'd expect in most classrooms or newsrooms, where the bulk of writing is exactly what it claims to be.

Feed 1,000 documents through that detector and watch what happens. Fifty of them are truly AI-written, and the tool catches around 47 or 48. The remaining 950 are truly human-written, yet even with 95% specificity, about 48 honest writers still get wrongly flagged as AI.

Add it up: roughly 48 false accusations against about 48 correct catches. Nearly half of every flag the detector produces is wrong, even though both its sensitivity and specificity individually look excellent. This is the base-rate fallacy at work, and it's the identical math behind the classic mammography example: a screening test can be 93%+ accurate on both ends and still mean most positive results are false alarms, purely because the condition being screened for is rare in the tested population (a documented illustration works out to roughly 91% of positive mammogram readings being false positives at real-world cancer prevalence and specificity rates).

The paper quantifies this specifically for one real, documented case: OpenAI's own AI-text classifier, which the company itself disclosed (January 31, 2023) "incorrectly labelled human-written text as AI-written 9% of the time, and only correctly identified 26% of AI-written texts." That's a 9% false-positive rate on its own, before prevalence even enters the picture, on a tool a major AI lab built and then withdrew for accuracy reasons six months later.

What this changes about reading a detection score

It is wrong to call detectors useless. The percentage listed in marketing materials reflects performance under a particular test setup, not the probability that any given flagged file truly originates from an AI. People often ask for the tool's overall accuracy, yet the relevant metric is the chance that a singled-out document is indeed AI-generated. That chance drops sharply as AI-generated material becomes rarer, widening the gap between the two measures.

That's exactly why TheChecker.AI's own reporting leans on a probability score instead of a pass/fail verdict. Our accuracy benchmark publishes the actual test data instead of one headline number, because a score only tells you how text compares statistically to known AI-generated patterns. It was never designed as a courtroom-grade verdict against one specific person, and this paper adds to a pattern this blog has tracked all year: a 99%-accurate detector still isn't proof of anything on its own, and research on false positives keeps landing on the same conclusion from a different angle.

Avoiding a single score for detection is the paper's suggested practical fix. Tools that detect misuse ought to combine a numeric score with a conversation, draft history, or a second independent check before any flag is considered definitive. Such caution proves vital in environments where real misuse is rare, classrooms, newsrooms, or peer-review desks.

When a detector with credible sensitivity and specificity evaluates every manuscript or application, and only a minority of those submissions conceal AI content, the detector will identify more false claims than genuine ones in sheer numbers. The mathematics predicts this result; it is not a rare or accidental outcome. Where the detector is applied changes its significance. A newsroom or admissions office that performs uniform scans across a low-risk cohort finds itself in a different situation than an investigator who scrutinizes a single suspect file against a flood of contradictory data. The rate of AI-generated content varies by context, and a sound process must adjust the influence of a raw score to reflect that variability.

FAQ

If a detector is 99% accurate, why would most of its flags still be wrong? The detector's 99% accuracy figure is misleading if you focus only on individual flags. That statistic reflects overall test performance, not the probability that any one flagged document is actually AI-written. That probability hinges on prevalence, the frequency of hidden AI output in the group under test. In situations where misuse is rare, a detector with very high accuracy will still produce more false positives than true positives when you simply count the flags.

Is this the base-rate fallacy from statistics? This text identifies the situation as an instance of the base-rate fallacy. It explains that the logic mirrors the familiar mammography example, in which a highly accurate test still yields predominantly false positives due to the low prevalence of the disease in the population screened. The 2026 Next Research paper transposes this framework to AI-text detectors. Furthermore, the paper extends Ioannidis's 2005 analysis that most research results are statistically prone to being false.

Does this mean AI detectors shouldn't be used at all? The study does not claim that AI detectors ought to be abandoned entirely. It contends that a single detection score should not determine the outcome on its own. This point gains particular importance when the incidence of AI-generated text is low. The highest utility of a detection score emerges when it is combined with process evidence. Process evidence may consist of a draft's revision sequence or a direct interview with the author. Consequently, a detection score must never trigger an automatic accusation.

See a probability-based score in practice

Understanding what a score can and can't prove is the first step to using one responsibly. Try TheChecker.AI's free demo to see how a probability-based, context-aware score reads in practice, instead of a single pass/fail number.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.

Interested in using TheChecker.AI?

Try it free