AI-written content now accounts for a growing share of the web, according to a Pew Research study published this week. The findings lend weight to the so-called Dead Internet Theory, a concept that emerged in 2021 and holds that much of what appears online today is generated by machines rather than people.
Researchers scanned nearly half a million English-language webpages published over the past five years, drawing on Common Crawl, a web archive. That timeframe stretches back to before ChatGPT became publicly available in November 2022. The pages were run through Open Pangram, an AI detection tool, to assess whether the text had been written or edited by artificial intelligence. Researchers then examined a sample of 10,000 pages from the previous month, looking for phrases, words and language patterns more commonly associated with AI systems than with human authors.
The team was clear that AI detection models are not infallible. They can flag human writing as machine-generated and miss text produced by AI. Even so, the study found that 10% of the samples showed significant signs of AI authorship.
One in 10 webpages as of July 2026 may appear modest, but the researchers noted that the internet contains a mixture of new and old material, and many pages in the random samples could not have been produced by AI. Filtering out older pages and focusing only on those published after ChatGPT’s release made the trend more pronounced. Researchers say the shift began in 2022 with ChatGPT and accelerated as other systems launched, including Claude and Google’s Gemini.
Where AI content appears most
The study shows that AI-derived text is most common on .com domains, at roughly 1 in every 10 pages. Other domains carried less machine-written material, with 4.6% for .org and around 1% for .edu and .gov webpages.
How to spot AI-written text
Certain punctuation, words and phrases turn up more often in AI-generated text than in human writing. Em dashes, for instance, appear twice as frequently, negative parallelisms are more common, and Oxford commas are used 63% more often than in human-written text. This largely reflects the fact that these systems are trained on human datasets and mimic that style, sometimes to excess. AI is also more likely to reach for words such as “delve” and “interplay”.
The researchers said detection models learn to recognise subtle statistical patterns in word choice and sentence structure, helping them judge whether a piece of writing was likely AI-authored. Other common signals include key terms from a prompt being repeated, repetitive explanations and a bot-like tone rather than a natural one. AI systems themselves often warn that their responses may be inaccurate, advising users to verify claims elsewhere online. According to the study, the sources used for that verification may increasingly carry AI-generated content of their own.