---
title: "ChatGPT Has Already Polluted the Internet So Badly That It’s Hobbling Future AI Development"
description: "There may be no undoing the vast amounts of pollution wreaked by ChatGPT. And that's just tough luck for any AI models that come after it."
date: "2025-06-16"
modified: "2025-06-16"
authors:
  - name: "Frank Landymore"
    job_title: "Contributing Writer"
    link: "https://futurism.com/authors/flandymore"
url: "https://futurism.com/chatgpt-polluted-ruined-ai-development"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai models"
  - "chatgpt"
  - "generative ai"
  - "OpenAI"
---

# ChatGPT Has Already Polluted the Internet So Badly That It’s Hobbling Future AI Development

![There may be no undoing the vast amounts of pollution wreaked by ChatGPT. And that's just tough luck for any AI models that come after it.](<https://futurism.com/wp-content/uploads/2025/06/chatgpt-polluted-ruined-ai-development.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

The rapid rise of ChatGPT — and the cavalcade of competitors' generative models that followed suit — has polluted the internet with so much useless slop that it's already kneecapping the development of future AI models.

As the AI-generated data clouds the human creations that these models are so heavily dependent on amalgamating, it becomes inevitable that a greater share of what these so-called intelligences learn from and imitate is itself an ersatz AI creation.

Repeat this process enough, and AI development begins to resemble a maximalist game of telephone in which not only is the quality of the content being produced diminished, resembling less and less what it's originally supposed to be replacing, but in which the participants [actively become stupider](<https://futurism.com/ai-industry-problem-smarter-hallucinating>). The industry likes to describe this scenario as AI "[model collapse](<https://futurism.com/the-byte/ai-trained-with-ai-generated-data-gibberish>)."

As a consequence, the finite amount of data predating ChatGPT's rise becomes extremely valuable. In a [new feature](<https://www.theregister.com/2025/06/15/ai_model_collapse_pollution/>), *The Register* likens this to the demand for "low-background steel," or steel that was produced before the detonation of the first nuclear bombs, starting in July 1945 with the US's Trinity test.

Just as the explosion of AI chatbots has irreversibly polluted the internet, so did the detonation of the atom bomb release radionuclides and other particulates that have seeped into virtually all steel produced thereafter. That makes modern metals unsuitable for use in some highly sensitive scientific and medical equipment. And so, what's old is new: a major source of low-background steel, even today, is WW1 and WW2 era battleships, including a huge naval fleet that was scuttled by German Admiral Ludwig von Reuter in 1919.

Maurice Chiodo, a research associate at the Centre for the Study of Existential Risk at the University of Cambridge called the admiral's actions the "greatest contribution to nuclear medicine in the world."

"That enabled us to have this almost infinite supply of low-background steel. If it weren't for that, we'd be kind of stuck," he told *The Register*. "So the analogy works here because you need something that happened before a certain date."

"But if you're collecting data before 2022 you're fairly confident that it has minimal, if any, contamination from generative AI," he added. "Everything before the date is 'safe, fine, clean,' everything after that is 'dirty.'"

In 2024, Chiodo co-authored a [paper](<https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5045155>) arguing that there needs to be a source of "clean" data not only to stave off model collapse, but to ensure fair competition between AI developers. Otherwise, the early pioneers of the tech, after ruining the internet for everyone else with their AI's refuse, would boast a massive advantage by being the only ones that benefited from a purer source of training data.

Whether model collapse, particularly as a result of contaminated data, is an imminent threat is a matter of some debate. But many researchers have been [sounding the alarm](<https://futurism.com/the-byte/ai-running-out-data-smarter>) for years now, including Chiodo.

"Now, it's not clear to what extent model collapse will be a problem, but if it is a problem, and we've contaminated this data environment, cleaning is going to be prohibitively expensive, probably impossible," he told *The Register*.

One area where the issue has already reared its head is with the technique called retrieval-augmented generation (RAG), which AI models use to supplement their dated training data with information pulled from the internet in real-time. But this new data isn't guaranteed to be free of AI tampering, and [some research](<https://futurism.com/ai-models-falling-apart>) has shown that this results in the chatbots producing far more "unsafe" responses.

The dilemma is also reflective of the broader debate around [scaling](<https://futurism.com/ai-researchers-tech-industry-dead-end>), or improving AI models by adding more data and processing power. After OpenAI and other developers reported diminishing returns with their [newest models in late 2024](<https://futurism.com/the-byte/openai-diminishing-returns>), some experts proclaimed that scaling had hit a "wall." And if that data is increasingly slop-laden, the wall would become that much more impassable.

Chiodo speculates that stronger regulations like labeling AI content could help "clean up" some of this pollution, but this would be difficult to enforce. In this regard, the AI industry, which has [cried foul](<https://futurism.com/the-byte/openai-eu-sam-altman>) at any [government interference](<https://futurism.com/openai-over-copyrighted-work>), may be its own worst enemy.

"Currently we are in a first phase of regulation where we are shying away a bit from regulation because we think we have to be innovative," Rupprecht Podszun, professor of civil and competition law at Heinrich Heine University Düsseldorf, who co-authored the 2024 paper with Chiodo, told *The Register*. "And this is very typical for whatever innovation we come up with. So AI is the big thing, let it go and fine."

**More on AI:** *[Sam Altman Says "Significant Fraction" of Earth's Total Electricity Should Go to Running AI](<https://futurism.com/openai-altman-electricity-ai>)*

## Author
At Futurism, my work has often centered on bringing a sense of clarity and insight to complex topics ranging from the regulation of emerging technologies to the esoteric ideologies of Silicon Valley executives, while striving not to lose the poetic sense of awe inspired by often-obscure fields like astrophysics and quantum computing. I broke the story of CNET using AI to produce articles that turned out to be riddled with factual errors and plagiarism — a dam-breaking inflection point, as I've reported, that's inspired copycats and endless discourse while beguiling stakeholders ranging from tech giants to purveyors of spam around the web. My work at Futurism has been cited by publications including CBS News, the Los Angeles Times, Vice, Gizmodo, Engadget, the Verge, and Vanity Fair. I grew up in locales ranging from India to China, and now live in the exotic suburbs of Virginia. In my free time, I'm an avid reader of weird sci-fi literature, an aficionado of East Asian cinema, and, regrettably, a relapsed gamer. Allegedly, I’m working on a debut novel, currently untitled.

### Author social links  
[Bluesky](<https://bsky.app/profile/f-w-l.bsky.social>)