The post-human web: 35% of new pages are the work of AI
A Pew Research Center report reveals how automation is transforming the fabric of the Internet and what it means for our perception of truth.
August 28, 2026 · 4 min read
TL;DR: A Pew Research Center study confirms that 35% of new pages since 2023 show signs of AI. Automation is reshaping the commercial web, prioritizing volume over human value.
The synthetic tide: a paradigm shift on the web
Since the emergence of ChatGPT in November 2022, the Internet's infrastructure has undergone a metamorphosis that redefines our relationship with information. An exhaustive study by the Pew Research Center suggests that we are not simply facing a technological evolution, but the potential 'death of the human web' as we knew it. The data is striking: more than 35% of pages published since the massive deployment of large language models (LLMs) show clear evidence of authorship assisted or generated by artificial intelligence. This phenomenon marks a milestone comparable to the transition from the static web to Web 2.0, but with a fundamental difference: the speed of adoption has been exponential, not linear.
The analysis of a sample of 10,000 pages extracted from the Common Crawl project reveals a telling historical contrast. While only 10% of the Internet's total history—a trajectory of more than 30 years—appears to have an artificial origin, the concentration of this content in the last two years is overwhelming. We are witnessing a structural shift where the metric of success has shifted from depth and veracity toward algorithmic efficiency. This model of "mass production" is reminiscent of the content farm bubble of 2010, but powered by a synthesis capability that scales without direct human intervention.
The digital footprints of automation
Identifying synthetic authorship is not an exact science, but the Pew Research Center report has managed to systematize recurring linguistic patterns that act as biological markers of AI. Beyond anecdotes, the study identifies syntactic structures that reveal the training architecture of models like GPT-4 or Claude. The excessive use of em dashes to connect ideas, the systematic increase in the use of the Oxford comma, and the repetition of binary structures like "it is not just X, it is Y" have become the fingerprints of current LLMs.
Furthermore, the lexicon has undergone a worrying standardization. Specific terms such as 'delve', 'interplay', and 'testament' have doubled in frequency in the commercial web compared to periods prior to 2022. This phenomenon is not accidental: it reflects the optimization of models for language that sounds "professional" but lacks the creative friction inherent to human thought. The proliferation of this synthetic content poses a direct threat to user trust, forcing us to question the origin and intent of every piece of data consumed on the web. It is, in essence, a challenge of authenticity reminiscent of the early days of the fight against spam, but where the enemy is not a rudimentary bot, but a system capable of simulating academic or journalistic prose with surgical precision.
Implications for the future of work and information
The Pew Research Center study highlights an alarming quality gap: .edu and .gov domains retain a higher proportion of human authorship, suggesting that institutions that prioritize verification and institutional accountability are the last barriers against automation. In contrast, the .com ecosystem, driven by advertising incentives and SEO, has been the most permeable, with AI rates significantly exceeding any other sector. For companies, this poses a challenge of reputation and survival: how to stand out in a sea of content generated by algorithms that, ironically, are designed to please other algorithms?
Speculation suggests that, in the near future, brands that manage to certify the human origin of their content (through digital signatures or blockchain technology) could command a trust premium. As the value of AI-generated content plummets due to saturation, the market could experience a return to the value of "human curation" and first-hand experience, assets that LLMs still cannot replicate without the risk of hallucinations or redundancy. The impact on the labor market is equally uncertain: the automation of technical and commercial writing lowers the barrier to entry, but at the same time devalues standard content production, forcing knowledge professionals to specialize in areas where AI still requires constant critical supervision.
What should readers know?
- Erosion of trust: The saturation of synthetic content makes fact-checking difficult, forcing users to seek primary sources and avoid aggregators.
- Commercial bias: Most AI is applied in marketing and monetization environments, where the goal is search engine positioning, not knowledge transfer.
- Evolution of critical reading: The modern reader must develop new skills, such as the ability to detect low-quality syntactic patterns and cross-verification using advanced search tools.
Although AI is a powerful tool for productivity, its indiscriminate use to populate the web threatens to turn the Internet into an ecosystem of constant 'noise'. The paradox is clear: as the web fills with content created by machines to be read by machines, the value of human, reflective, and verifiable information becomes a scarce and, therefore, extremely valuable commodity. This is the challenge that search engines and curation platforms will have to solve in the coming years by implementing new indexing standards that penalize low-quality production and reward verifiable originality.