Back to Blog
Research7 min read

AI Generated Text Examples: What the Fingerprints Actually Look Like

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Layered paper-cut illustration of a torn-paper magnifying glass held over indigo ink-wash manuscript pages, an amber thread tracing a jagged sentence-length graph, representing measurable AI-text fingerprints

Quick answer

"AI-generated text" isn't one style. It's a set of measurable statistical habits that show up across models, and researchers have now quantified exactly which words and sentence patterns give it away. This isn't a vibe check. Peer-reviewed studies have counted the actual excess words, measured the actual sentence-length variance, and published the actual numbers. Below are real, sourced examples of what those fingerprints look like in practice, and what a detector is actually measuring when it flags them.

The vocabulary fingerprint: real words, real frequency data

The clearest, most measured example of an AI-text fingerprint comes from a 2024 study by Dmitry Kobak and colleagues at the University of Tübingen, who analyzed over 14 million PubMed biomedical abstracts published between 2010 and 2024 (Kobak et al., "Delving into ChatGPT usage in academic writing through excess vocabulary," arXiv:2406.07016). They compared actual 2024 word frequencies against a counterfactual: what those frequencies would have been if pre-ChatGPT (2021-2022) trends had simply continued. Words that spiked far beyond that projection are what they called "excess vocabulary," and the list is specific and reproducible.

"Delves" (and its inflections "delve," "delved," "delving") tops the list, nearly 25 times more frequent than expected for 2024. "Showcasing" trails close behind at a ratio near 9.2. "Underscores" isn't far off, around 9.1. Ordinary words moved too. "Potential" showed up in 4 more percentage points of abstracts than the pre-ChatGPT trend predicted, and "crucial" and "findings" climbed as well. Here's what that flowery, AI-preferred phrasing reads like, pulled straight from the paper's own examples:

"By meticulously delving into the intricate web... takes a deep dive into their involvement as significant..."

"A comprehensive grasp of the intricate interplay between [...] is pivotal for effective therapeutic strategies."

Those lines aren't invented for this article. They're quoted directly from the study as real illustrations of a pattern researchers found at scale. The team's bottom line: at least 10% of all 2024 PubMed abstracts show signs of LLM involvement, and in fields like computer science and bioinformatics, that lower-bound estimate climbs past 20%.

A separate, more recent study replicated the same core finding in medical writing specifically. Kentaro Matsui's 2025 analysis in Perspectives on Medical Education tracked 135 potentially AI-influenced terms across PubMed records from 2000-2024 against 84 control phrases common in medical writing (Matsui, "Delving Into PubMed Records," PMC12679996). 103 of those 135 terms showed statistically meaningful frequency increases by 2024, with "delve," "underscore," "primarily," "meticulous," and "boast" topping the list. One nuance worth keeping honest: Matsui's data shows these terms were already trending upward starting around 2020, two years before ChatGPT existed, so the pattern isn't purely an AI artifact. It's a pre-existing trend generative AI sharply accelerated rather than invented from scratch.

The structural fingerprint: sentences, punctuation, and paragraph shape

Vocabulary is only half of it. Researchers at the University of Kansas built a classifier in 2023 that separated human academic writing from ChatGPT output with over 99% document-level accuracy. It didn't read for meaning at all. It counted 20 structural features (Desaire et al., "Distinguishing academic science writing from humans or ChatGPT," PMC10328544). Here's what actually mattered:

  • Sentences and words per paragraph. Human scientists wrote longer paragraphs with more sentences. ChatGPT's answers ran shorter, tighter, more compact.
  • Standard deviation in sentence length. This is burstiness, measured as one number. Humans swung between short and long sentences constantly; ChatGPT stayed in a narrower band. People also reached for very short sentences (10 words or fewer) and very long ones (35+) far more than the model did, the two extremes AI text tends to avoid.
  • Punctuation habits. Humans reached for question marks, dashes, parentheses, semicolons, colons. ChatGPT leaned on single quotation marks instead.
  • Word choice for attribution. ChatGPT defaulted to vague collective phrasing, "others," "researchers." Human scientists named names.
  • Equivocal language. Humans used "however," "but," "although," and "because" more than the model did. That cuts against the usual assumption that AI text hedges more. Here it was the humans hedging more, just with different words than the generic AI hedges people expect ("it's important to note," "this can vary").

At the paragraph level alone, this feature set classified writing correctly 94% of the time. Combining all paragraphs in a document pushed accuracy to 99.5%, with only one document out of 192 misclassified, and that single miss was, tellingly, an obituary-style piece about a deceased scientist rather than a standard research summary, an edge case the model wasn't built to expect.

A worked example: same idea, two different statistical signatures

Put the two studies together and a pattern emerges that's genuinely reproducible, not anecdotal. Take a single idea, "this treatment shows promise for future patients," and notice how the two fingerprints show up differently depending on who wrote it:

An AI-typical rendering, matching the Kobak et al. and Desaire et al. patterns: "A comprehensive evaluation of the treatment's potential underscores its pivotal role in shaping future therapeutic approaches, with researchers noting promising preliminary findings." Notice the excess-vocabulary words (comprehensive, potential, underscores, pivotal, researchers, findings), the vague attribution ("researchers" instead of a name), and the single long, evenly-paced sentence carrying the whole idea.

A human-typical rendering of the same idea, matching the same research: "Dr. Alvarez's team saw something worth watching: patients who started treatment early did better. It's not proof yet, the sample was small, but it's the kind of early signal that makes you want to run the bigger trial." Notice the named attribution, the short punchy sentence followed by a longer qualified one (real burstiness), and the equivocal "but" doing real hedging work instead of a stacked wall of qualifiers.

Neither sentence is a gotcha. These are statistical tendencies, not rules. Plenty of human writers can produce a smooth, evenly-paced sentence loaded with "potential" and "underscores" with zero AI involvement. That's exactly why a real detector never grades one sentence alone. It scores the whole passage, the same way perplexity and burstiness get measured across a full document instead of a single line.

Why a handful of examples isn't the same as a verdict

These are real, sourced, quantified patterns, not folklore. But every one of them describes a tendency, and tendencies have exceptions in both directions: careful human writers sometimes produce evenly-paced, vocabulary-heavy prose, and increasingly capable models are trained to avoid the most obvious tells on purpose. That's the same honest caveat that applies to any list of AI-writing signs. A pattern match is a strong lead, not proof, which is exactly why detectors themselves can occasionally misfire on both sides: flagging real human writing that happens to share the pattern, or missing AI writing that's been edited to avoid it.

FAQ

Are "delve" and "underscore" proof that something was written by AI? No. Since 2022, AI-influenced writing has shown a statistically significant over-representation of these specific words. That's just one of many useful signals, not proof on its own. Human writers also commonly use those words, and with the pattern now widely recognized, more writers actively avoid them.

Do these examples apply outside of academic and medical writing? The two vocabulary studies above specifically analyzed PubMed abstracts, so the exact words and ratios are strongest evidence in academic and scientific writing. The structural patterns (paragraph length, sentence-length variance, hedging habits) are more general statistical tendencies that detection tools apply across genres, not just academic prose.

What's the actual difference between spotting these patterns by eye and running a real detection check? Reading for a handful of buzzwords catches only the most obvious cases and misses everything else a full statistical scan measures: sentence-level perplexity, paragraph-level burstiness, and cross-referencing against signatures from more than 40 AI systems. A worked example in an article is illustrative. A real score on your actual text is evidence.

Curious what these patterns look like when it's your own writing under the microscope? Paste it into TheChecker.AI's free demo and see the sentence-by-sentence breakdown instead of scanning for buzzwords by eye. For the statistics behind the score, see how AI text detection works, and if you've ever been on the wrong end of a false flag, read what the research says about when detectors get it wrong.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how educators and teams can use it responsibly.