More than a third of English-language web pages published since ChatGPT's late-2022 launch carry statistical fingerprints of AI authorship, according to a new Pew Research Center study that scanned roughly 490,000 pages pulled from the Common Crawl archive between January 2021 and July 2026.
Pew's researchers used Open Pangram, a detection model that flags AI involvement by analyzing statistical patterns in sentence structure and word choice rather than searching for specific telltale phrases. Across the full dataset spanning 2021 to 2026, roughly 10% of pages showed significant signs of AI authorship. Narrowed to only the pages published after ChatGPT's public debut, that figure jumps to 35%.
Commercial Web Pages Lead the Shift
The trend is heavily concentrated on commercial domains. Pages on .com sites showed AI-authorship signals at roughly 10 times the rate found on .edu and .gov domains, which stayed near 1% even as ChatGPT usage exploded. Pew also tracked a set of stylistic markers that have become associated with AI-generated prose: em-dash usage roughly doubled since 2023, Oxford comma use rose 63%, and formal vocabulary choices like "delve" and "testament" more than doubled in frequency, alongside a near-tripling of a pattern researchers call "negative parallelism" — sentence constructions like "it's not just X, it's Y."
Pew was careful to note the limits of its own method: the study's own methodology notes that detection models can misclassify pages in both directions, and that "significant signs of AI authorship" indicates likely AI assistance rather than proof a page was machine-written end to end.
Related: Microsoft Patches Maximum-Severity Entra ID Flaw Rated CVSS 10.0
Why This Matters Beyond SEO Spam
For the crypto industry specifically, the growth of AI-generated web content intersects with a less abstract problem: AI-assisted fraud. TRM Labs has reported a roughly 40% rise in AI-driven crypto crime, with deepfake-enabled scams climbing 13-fold compared to 2022 as fraudsters use synthetic audio, video and text to impersonate exchanges, executives and even blockchain investigators to push fake token giveaways and phishing links.
That backdrop is part of why blockchain-based content provenance has moved from a niche cryptography topic to an active product category. Systems built around standards like the Coalition for Content Provenance and Authenticity aim to cryptographically sign and timestamp media at the point of creation, giving readers a way to verify whether a video, image or article actually originated from the account it claims to — the same on-chain verification logic crypto platforms already use to trace stolen funds now being repurposed to trace fabricated content back to its source.
An Arms Race With No Finish Line
Pew's data suggests the AI-authorship trend is still accelerating rather than plateauing, which means the gap between how much content looks human-made and how much actually is will likely keep widening. For an industry already grappling with deepfake-driven scams impersonating everyone from exchange executives to on-chain investigators, a web where a third of new writing carries AI fingerprints raises the stakes on provenance tools that can tell readers, and investors, what to actually trust.