Back to Blog
Research & Data 8 min read

Pew Says a Third of New Web Pages Are AI-Written. Here's What That Actually Means

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Paper-cut diorama of a globe made of layered web-page sheets, a portion tinted with an ink-wash grid pattern to represent AI-authored content, beside an abstract circular seal

Quick answer

Pew Research Center's July 2026 analysis examined 490,000 English-language web pages pulled from Common Crawl. Detectable AI-authorship signals appeared on one in ten of those pages. If the analysis is narrowed to sites that appeared after the November 2022 debut of ChatGPT, that share rises to slightly over a third. Stanford, Imperial College London, and the Internet Archive conducted related research with a different archive and detection tool. They reported that about 35% of new sites by mid-2025 exhibited AI-authorship signals. Two distinct research groups applied separate methodologies, yet the core lesson remains: the findings align, not clash.

What Pew actually measured

No survey or self-reporting from AI companies was conducted by Pew. From January 2021 to July 2026, researchers extracted 10,000 random English-language pages from each of 49 Common Crawl snapshots. In all, 490,000 pages were examined. An open-weight AI detection model from Pangram processed the body text of each page. Scores ranged from 0, indicating fully human, to 1, indicating fully AI. Pages scoring 0.2 or higher were counted as exhibiting "meaningful" AI authorship or editing.

Pangram's flagship commercial model processed 62,370 of those pages. It agreed with the open-weight model's classification on 96% of them. The calculation gives a Cohen's kappa of 0.61, a value that signals strong, yet imperfect, concordance. A dataset of half a million pages reveals numerous edge-case disagreements, and each one matters. Earlier scans of open-model material found roughly a one percent false-positive rate on overtly human-written text. The score remains a probability estimate, not an absolute judgment on any single page or paragraph.

Where the AI writing actually concentrates

At first, the web's reporting on the story seems spotty, but the underlying numbers tell a different story. In 2026, about 9.35 percent of .com pages carry AI-authorship signals. .org sites trail at approximately 4.59 percent, which is almost half the rate of .com. Educational domains drop to 1.03 percent, roughly one-tenth of the .com figure. Government websites register only 0.76 percent, also near one-tenth of .com. This disparity did not exist in 2021, when the four domain types were within one percentage point of each other. The divergence began after November 2022 and has widened since.

Fast adoption of AI tools in commercial content mills, affiliate sites, and SEO-driven blogs stems from the need to produce high volumes. Government and university pages, however, advance at a slower pace, governed by institutional review and a lower drive for volume. The domain-level distinction primarily signals different publishing motivations, not that .com material is inherently poorer.

What the AI "tells" actually look like right now

Pew's team conducted a controlled study. The study searched for linguistic markers within a consistent sample. They collected 2026 web text. They compared this data with the 2023 baseline. The comparative analysis exposed quantifiable changes in phrasing and terminology.

Writers now employ em dashes about twice as frequently as before. Oxford commas have risen by 63 percent. Researchers have singled out a particular cluster of words AI models tend to repeat, Pew's list highlights terms such as "delve," "interplay," "testament," "underscore," "pivotal," "meticulous," and others, and this cluster's frequency has more than doubled. The "negative parallelism" construction, the "it's not just X, it's Y" pattern, has increased to nearly three times its earlier rate, though it remains uncommon in the wider sample.

Origin cannot be pinned to one marker. Those who craft text with painstaking care use em dashes. The usefulness of a signal appears only when many markers are examined together. Pew insists that detection rests on the collective signal: a reliable detection model examines a wide set of statistical patterns at once, not just a single identifying phrase. For a deeper look at what those fingerprints look like inside individual documents, with quoted examples from peer-reviewed studies, see our earlier piece on AI-generated text fingerprints.

A second study, a different method, the same rough number

In a partnership overseen by Jonas Dolezal, Stanford, Imperial College London, and the Internet Archive joined forces. Their aim was to ask the same research question, but from another angle. Data collection switched from Common Crawl to the Wayback Machine. The sampling scheme was stratified, designed to create a uniform cross-section of the public web, a departure from Pew's snapshot approach. Testing several detectors against the RAID benchmark led the team to choose Pangram's flagship model. Their estimate: by mid-2025, about 35 percent of newly launched sites showed evidence of AI generation or assistance, a marked increase from the pre-ChatGPT era. Neither team had seen the other's numbers before publishing. The study's archive, sampling method, and detection model were all distinct, yet the one-third figure stayed consistent. The takeaway is that consistency across methods carries more weight than the precise number.

Researchers examined six feared impacts of AI-generated content on the web. These concerns involved a narrowing of vocabulary, a dip in factual precision, and a movement toward more uniform writing styles. Results supported just two of those worries. Specifically, the study did not find clear evidence that AI text lowers factual accuracy, nor that it erodes overall stylistic diversity. Yet the common assumption is that AI writing is making content less factual or stylistically diverse. The difference between what people think and what the data show is significant before any definitive judgments are made.

Academic and scientific writing wrestles with a problem similar to the one on the open web, though the underlying causes differ. Incentives that drive paper submission and publication aren't the same as the ones on open online platforms. If that specific angle interests you, our piece on AI writing in science papers covers a separate 12,750-paper arXiv study with its own methodology.

What a population-level number can't tell you

Citing "10% of the web" or "35% of new pages" describes the proportion of a larger population. It doesn't specify which exact pages those are. Pew's own writeup makes the point directly: an AI detection model is a probabilistic tool, and a classification on a single page isn't proof of who wrote it. It's an estimate, not a final judgment.

A detection score alone can't conclusively verify a document, essay, or cover letter. Even an accurate score only indicates a probability, not certainty. Human judgment remains indispensable, the score can't replace it, and treating a single classification as a "gotcha" misreads the tool's purpose. See why a 99% accurate detector still isn't proof of anything for the fuller version of that argument.

Why this matters beyond "the web has more AI text now"

This applies to a cover letter too, and to the portfolio page it links to, and to whatever "About" page a hiring manager clicks next — see our breakdown of what hiring managers can and can't actually tell about AI-written applications. All of that lives on the same web these numbers describe. Once a third of competitor content carries AI assistance, undisclosed use stops being theoretical for a content team. It's already sitting in the search results next to yours. A quote pulled off a random page now carries a real chance of being AI-written. Not automatically wrong. Just worth a second look.

Within the last four years, AI-authored text has moved from a novelty to a measurable part of new publishing. That growth is seen where content volume pays the bills. Knowing the aggregate rate of AI-authored text is useful. However, deciding whether a specific page, application, or document is AI-written is a different question. A real detector answers the question of AI authorship far better than a guess. If you evaluate content, applications, or a site's authenticity, check it for free rather than assume.

FAQ

Is Pew's study accurate? Pew double-checked its own results by running a subset of pages through a second, more advanced Pangram model, and got a 96% match between the two systems. That's a solid consistency check. Pew is upfront about the limit, though: any one page's classification stays a probability, never a guarantee.

Does this mean a third of everything I read online is fake or wrong? In the Dolezal et al. study, researchers checked whether rising AI authorship online lined up with a drop in factual accuracy. They found no statistically significant link. An AI-generated passage can still be correct, and human-written text can still be wrong. Authorship and accuracy are separate questions.

Why is AI writing so much more common on .com domains than .edu or .gov sites? Pew's data show .com domains carrying AI-authorship signals at roughly ten times the rate of .edu or .gov domains. The simplest explanation is publishing incentives. Commercial and content-driven sites benefit from high-volume output in a way institutions and government agencies generally don't, so they picked up AI writing tools faster.

Can I check whether a specific page, document, or piece of writing was AI-generated? Ever wonder if a particular page, document, or piece of writing is AI-generated? A real-time detector was built exactly for that. Population studies like Pew's chart trends across hundreds of thousands of pages.

What detection method did Pew use? Open Pangram, an open-weight detection model, was applied to each of the 490,000 pages sampled. Consistent use of that model across the entire dataset kept the analysis coherent. For further confidence, a smaller portion of the data was run through Pangram's more sophisticated commercial model. That smaller test confirmed the pattern.

Check your own page before you publish it

Within the last four years, AI-authored text has moved from a novelty to a measurable part of new publishing, concentrated where content volume pays the bills. Knowing the aggregate rate is useful. Deciding whether a specific page, application, or document in front of you was AI-written is a different question, and a real detector answers that far better than a guess. If you're evaluating content, applications, or your own site's authenticity right now, check it for free rather than assume.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.

Interested in using TheChecker.AI?

Try it free