---
title: "AI Models Show Signs of Falling Apart as They Ingest More AI-Generated Data"
description: "As CEOs trip over themselves to invest in AI, the models are falling apart at the seams and going mad from cannibalism."
date: "2025-05-30"
modified: "2025-05-30"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/ai-models-falling-apart"
categories:
  - "Artificial Intelligence"
tags:
  - "ai training"
  - "llms"
  - "retrieval-augmented generation"
  - "synthetic data"
---

# AI Models Show Signs of Falling Apart as They Ingest More AI-Generated Data

![As CEOs trip over themselves to invest in AI, the models are falling apart at the seams and going mad from cannibalism.](<https://futurism.com/wp-content/uploads/2025/05/ai-models-falling-apart.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

As [CEOs trip over themselves](<https://futurism.com/ceos-return-ai-investments>) to invest in artificial intelligence, there's a massive and growing elephant in the room: that any models trained on web data from after the advent of ChatGPT in 2022 are ingesting AI-generated data — an act of low-key cannibalism that may well be causing increasing technical issues that could come to threaten the entire industry.

In a [new essay for *The Register*](<https://www.theregister.com/2025/05/27/opinion_column_ai_model_collapse/>), veteran tech columnist Steven Vaughn-Nichols warns that even attempts to head off so-called "model collapse" — which occurs when large language models (LLMs) are fed [synthetic, AI-generated data](<https://futurism.com/the-byte/what-is-synthetic-data>) and consequently [go off the rails](<https://futurism.com/the-byte/ai-trained-with-ai-generated-data-gibberish>) — are another kind of nightmare.

As [*Futurism*](<https://futurism.com/the-byte/ai-dumber>) and [countless](<https://www.bbc.com/audio/play/m00274wj>) [other](<https://www.nytimes.com/interactive/2024/08/26/upshot/ai-synthetic-data.html>) [outlets](<https://apnews.com/article/ai-artificial-intelligence-training-data-running-out-9676145bac0d30ecce1513c20561b87d>) have reported over the [last few years](<https://futurism.com/ai-trained-ai-generated-data-interview>), the AI industry has continuously barreled toward the moment at which all available authentic training data — that is, information that was produced by humans and not AI — [will be exhausted](<https://futurism.com/the-byte/ai-training-data-shortage>). Some pundits, [including Elon Musk](<https://techcrunch.com/2025/01/08/elon-musk-agrees-that-weve-exhausted-ai-training-data/>), believe we're already there.

To circumvent this "[Garbage In/Garbage Out](<https://www.technologyreview.com/2024/07/24/1095263/ai-that-feeds-on-a-diet-of-ai-garbage-ends-up-spitting-out-nonsense/>)" conundrum, industry titans including Google, OpenAI, and Anthropic have engaged in what's known as retrieval-augmented generation (RAG), which essentially involves plugging LLMs up to the internet so they can look things up if they're presented with prompts that don't have answers in their training data.

That concept seems pretty intuitive on its face, especially when presented with the specter of rapidly-approaching model collapse. There's only one problem: the internet is now full of lazy content that uses AI to drum up answers to common questions, often with hilariously bad and inaccurate results.

In a recent study from the research arm of Michael Bloomberg's media empire that was [presented at a computational linguistics conference](<https://aclanthology.org/2025.naacl-long.281/>) in April, 11 of the latest LLMs, including OpenAI's GPT-4o, Anthropic's Claude-3.5-Sonnet, and Google's Gemma-7B, produced far more "unsafe" responses than their non-RAG counterparts. As the paper put it, those safety concerns can include "harmful, illegal, offensive, and unethical content, such as spreading misinformation and jeopardizing personal safety and privacy."

"This counterintuitive finding has far-reaching implications given how ubiquitously RAG is used in \[generative AI\] applications such as customer support agents and question-answering systems," explained Amanda Stent, Bloomberg's head of AI research and strategy, in another interview with Vaughn-Nichols [published in *ZDNet*](<https://www.zdnet.com/article/rag-can-make-ai-models-riskier-and-less-reliable-new-research-shows/>) earlier this month. "The average internet user interacts with RAG-based systems daily. AI practitioners need to be thoughtful about how to use RAG responsibly."

So if AI is going to run out of training data — or it has already — and plugging it up to the internet doesn't work because the internet is now full of AI slop, where do we go from here? Vaughn-Nichols notes that [some folks](<https://venturebeat.com/ai/synthetic-data-has-its-limits-why-human-sourced-data-can-help-prevent-ai-model-collapse/>) have [suggested mixing authentic and synthetic](<https://arxiv.org/abs/2404.01413>) to produce a heady cocktail of good AI training data — but that would require humans to keep creating real content for training data, and the AI industry is actively undermining the incentive structures fo them to continue — while [pilfering their work without permission](<https://futurism.com/nick-clegg-scoffs-ai-copyright>), of course.

A third option, Vaughn-Nichols predicts, appears to already be in motion.

"We're going to invest more and more in AI, right up to the point that model collapse hits hard and AI answers are so bad even a brain-dead CEO can't ignore it," he wrote.

**More on AI in crisis:** [*Legendary Facebook Exec Scoffs, Says AI Could Never Be Profitable If Tech Companies Had to Ask for Artists' Consent to Ingest Their Work*](<https://futurism.com/nick-clegg-scoffs-ai-copyright>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)