AI Detection Score Meaning: What That Percentage Actually Tells You
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Quick answer
A number like "62% AI" is not a measurement of how much of the text is AI. It is a probability — how confident the model is that the whole document was machine-generated, based on patterns it learned from millions of human and AI samples. Every major detector, including ours, works this way. A score is a signal to investigate further, never a verdict on its own.
Why this confuses almost everyone the first time
Someone gets a report back that says "18% AI" and reads it as "18% of my essay was written by AI." That is not what the number means. Detectors compare the statistical fingerprint of your whole document — word choice, sentence-length variation, predictability — against patterns typical of human writing versus machine writing, then output a probability that the document as a whole came from an AI model.
GPTZero's own help documentation states this directly: "the percent you're seeing is NOT saying that '6% of your document contains AI.' GPTZero's percentage refers to our model's probability that your document was written by either AI or human." GPTZero even breaks it down further — a "4% AI" result, per their own example, means the model believes roughly 4 times out of 100 similar cases the document would be fully AI, 57 times out of 100 a mix of human and AI text, and 39 times out of 100 fully human. It's a distribution, not a slice.
Turnitin's documentation makes a related but distinct point: their overall percentage reflects how much "qualifying text" (prose sentences) the model flags as likely AI-generated or AI-paraphrased — which is closer to a proportion of flagged content, but still explicitly hedged. Turnitin's own guides state the tool "may not always be accurate... so it should not be used as the sole basis for adverse actions against a student." Two different detectors, two related but not identical definitions of what the number represents — which is itself a reason no single score should be treated as a courtroom verdict.
Score bands: what they usually mean
A universal standard is absent across all detectors. Yet the majority of tools that release threshold data cluster around a common figure. This alignment means that detectors and tools share comparable thresholds.
When a model detects a very low pattern match, near zero up to twenty percent, it sees barely any AI-typical signal. Because of this, Turnitin began showing an asterisk instead of a percentage for any score below twenty percent starting in July 2024, after its own testing found a higher incidence of false positives across that band. GPTZero refers to the same low range as “highly confident” human writing, with a stated error rate under two percent. The ambiguous middle range, roughly twenty to seventy percent, carries more uncertainty. Human writing that has been edited or polished with AI tools tends to land in that mixed category, and it's where GPTZero's own worked example (57 times out of 100 a mix of human and AI text) sits. A score in this range should prompt a closer look, not an accusation. At the high end, roughly seventy to a hundred percent, the pattern strongly points to AI generation. Still, Turnitin's own writing-detection guide states that even a fully-processed, non-asterisked score “should not be used as the sole basis for adverse actions.”
Turnitin's Chief Product Officer, Annie Chechitelli, put a number on the company's own real-world catch rate in 2023: roughly 85% of AI writing detected, alongside a deliberate trade-off to reduce false accusations. That's the company's own estimate of its accuracy, not an independent third-party audit — worth knowing when someone treats a Turnitin score as unimpeachable. See our own published accuracy benchmarks for how we test and report ours.
Confidence is a second, separate number
GPTZero's documentation draws a distinction that's easy to miss: the score (a probability) and the confidence level (how reliable that specific prediction is) are two different things. GPTZero publishes three confidence bands: "highly confident" (error rate under 2%), "moderately confident" (around 10%), and "low confidence" (14% or higher). A document can return a mid-range AI probability with low confidence attached — which is the software's own way of saying "don't lean on this one." Most detector interfaces show this pairing somewhere in the report; it's worth reading past the headline percentage to find it.
Why a raw percentage still isn't enough on its own
Academic-integrity offices increasingly say this out loud. The University of Melbourne's public guidance for students states plainly: "an AI writing detection report alone is not sufficient evidence for an allegation" of misconduct, and being asked to explain a flagged submission "is not an allegation of academic misconduct" — it is meant to be an informal, exploratory conversation. That framing matters because a score, by itself, describes a pattern match, not an observed act.
There's also a structural reason scores drift with the writer. Research on detector false-positive rates has repeatedly found that writing which is unusually short, formulaic, or highly polished — the kind non-native English speakers are often taught to produce — scores as more "AI-like" even when it is fully human. That's not a flaw specific to one vendor; it's a consequence of training detectors on the statistical gap between typical human variation and typical machine consistency, and formulaic human writing sits closer to the machine side of that gap. A score means less when the writer's style naturally resembles the pattern the model is watching for — see our own breakdown of how detectors can be wrong for the research behind this.
What to actually do with a score
Treat the percentage as a probability, not a fraction of the text. Check the report's own words, confidence level notes, sentence-by-sentence highlights, first. Only then act on the figures. Don't act on a single low-confidence or borderline result alone. Ask for a second opinion, a different detector, a conversation about the writing's development, or a version history, before treating a mid-range result as settled. Run your own check before submitting anything you're unsure about. TheChecker.AI's free demo hands back a detailed sentence-level breakdown, not one overall number, so you can spot exactly which parts of a document are pulling the score up or down before anyone else does. If you manage a policy (teacher, editor, hiring manager), write the threshold down. Set in advance what score band leads to a conversation, what maps to an audit, and what falls below the line, and say so publicly. Ambiguity is where false-positive damage happens.
FAQ
Is a 20% AI score bad? Not by itself. Turnitin's own published policy, along with most other detectors, treats sub-20% scores as statistically unreliable rather than clean evidence of anything. Because that range produces more false positives, Turnitin no longer shows an exact number below the line.
Why did my writing get a mixed score when I wrote it myself? A rigid template governs the writing. The sentences are short. A grammar-assistant gets applied heavily. The final rating leans toward mixed despite zero AI involvement. This outcome exposes a known limit of pattern-based detection. It offers no evidence that anything went wrong.
Can two detectors give different scores for the same text? Yes, routinely. Different models are trained on different data and weight different signals (sentence burstiness, vocabulary predictability, structural regularity), so disagreement between tools is expected, not a bug in either one. Read more on how detectors are built.
Does a high score mean someone definitely used AI? No. A strong pattern match raises a question, not a verdict. Whether it's enough to act on depends on the stakes, the confidence level attached to the score, and whether there's supporting evidence beyond the number itself.
---
Want to understand how each line holds up? Run your own writing through TheChecker.AI's [free demo](/demo) and get a sentence-level breakdown. The tool offers the same nuanced insight you need, but instead of a single number, it shows how your text performs one sentence at a time.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how educators and teams can use it responsibly.
Related posts

Is AI Detection Accurate? What the 2026 Research Actually Shows
Two new 2026 studies and the field's 2023 baseline, compared: false positives are down sharply, but catching fully AI-generated text still varies wildly by tool, and hybrid text defeats almost everyone.
Read more
Does 'Humanizing' AI Text Actually Beat Detection? What the Research Shows
'Humanizer' tools promise to make AI writing undetectable. Peer-reviewed research and the detectors' own published data tell a messier story — here's what actually happens when you run humanized text through a real detector.
Read more
Can Turnitin Detect ChatGPT? What Turnitin's Own Data Actually Says
Yes, most of the time — but Turnitin's own published numbers show a real gap between its accuracy claim and what it admits about false positives. Here's what the company's data says, what two universities found when they tested it themselves, and what that means for a single score.
Read more