---
title: "When AI Is Trained With AI-Generated Data, It Starts Spouting Gibberish"
description: "What happens when you train an AI model with AI-generated content? Absolute chaos, according to a new study."
date: "2024-07-26"
modified: "2024-07-26"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/the-byte/ai-trained-with-ai-generated-data-gibberish"
categories:
  - "Artificial Intelligence"
tags:
  - "ai"
  - "ai training"
  - "ai-generated content"
  - "the digest"
---

# When AI Is Trained With AI-Generated Data, It Starts Spouting Gibberish

![What happens when you train an AI model with AI-generated content? Absolute chaos, according to a new study.](<https://futurism.com/wp-content/uploads/2024/07/ai-trained-with-ai-generated-data-gibberish.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

What happens when you feed AI-generated content back into an AI model? Put simply: absolute chaos.

A fascinating [new study](<https://www.nature.com/articles/s41586-024-07566-y>) published in the journal *Nature* shows that AI models trained on AI-generated material will experience rapid "model collapse." Basically, as an AI model cannibalizes AI-generated data, its outputs become increasingly bizarre, garbled, and nonsensical, as if synthetic data — as opposed to high-quality, human-made material — breaks its brain.

On the one hand, the study's results serve as [another reminder](<https://futurism.com/ai-trained-ai-generated-data-interview>) that AI models are incredibly responsive to their training data, and that allowing AI-generated material to seep into those datasets can have serious consequences for AI systems and the billion-dollar companies building them. At the same time, it underscores AI companies' [ever-growing need](<https://futurism.com/the-byte/ai-training-data-shortage>) for high-quality human material with which to train its models — an increasingly scarce, and thus increasingly valuable, resource that could stand to put generative AI advancement at a plateau.

"The message is we have to be very careful about what ends up in our training data," study co-author Zakhar Shumaylov, an AI researcher at the University of Cambridge, [told *Nature*](<https://www.nature.com/articles/d41586-024-02420-7>), warning that otherwise "things will always, provably, go wrong."

Shumaylov's team used a pre-trained large language model (LLM) which they then calibrated with a HuggingFace dataset comprised of Wikipedia entries. The researchers then put the model through a string of generations, each time returning the AI's output back into the training set.

The results were striking. A prompt about buildings in Somerset, England — the text of which was taken from [this niche Wikipedia page](<https://en.wikipedia.org/wiki/Wikipedia:Featured_topics/Grade_I_listed_buildings_in_Somerset>) — for example, first returned a relatively normal, though still error-stricken, response. But by the researchers' ninth iteration, the model's response was total gibberish about... jackrabbit tails.

"architecture," read the AI's garbled output. "In addition to being home to some of the world's largest populations of black @-@ tailed jackrabbits, white @-@ tailed jackrabbits, blue @-@ tailed jackrabbits, red @-@ tailed jackrabbits, yellow @"

However bizarre the results, the process of model collapse is actually fairly simple. An AI system only has access to the data that it's provided; more original, human-made data generally means a better-functioning generative AI system, as does diversity within that data. Conversely, feeding a model with AI-spun generations is diversity-limiting. The model will compound its own errors, forget certain words and artifacts that are less present in its training, and eventually cave in on itself.

The study's authors aren't the first to measure this phenomenon. AI researcher Jathan Sadowski last year dubbed the destructive process as "Habsburg AI," wherein an AI model fed AI-made content essentially becomes an "inbred mutant," much like Europe's infamously-inbreeding Habsburg family inbred itself into [infertility and decline](<https://www.smithsonianmag.com/smart-news/distinctive-habsburg-jaw-was-likely-result-royal-familys-inbreeding-180973688/>). Indeed, similarly to how humans need genetic diversity in reproduction to avoid historically recessive jawlines, an AI model seems to need high-quality diversity in its training data to avoid collapse.

The study also raises another serious point of concern for data-desperate AI companies, which is the wavering sustainability of web scraping. AI models have largely been trained on data that was scraped from the open web and social media. Now, though, the internet is increasingly chock-full of AI-generated content. There are thousands of [AI-powered, spammy "news" sites](<https://www.404media.co/i-paid-365-63-to-replace-404-media-with-ai/>) cropping up in Google; Facebook is quickly filling with bizarre AI imagery of [soldiers and Jesus](<https://futurism.com/the-byte/ai-facebook-pandering>); established media companies have, in a growing number of cases, published [AI-generated content](<https://futurism.com/advon-ai-content>) on their websites. Very little of this content is marked as AI-generated, meaning that web scraping, should AI companies continue to attempt to gather their data from the digital wilds, is becoming a progressively dubious means of collecting AI training data.

"The need to distinguish data generated by LLMs from other data raises questions about the provenance of content that is crawled from the Internet," the study authors write, adding that it's "unclear how content generated by LLMs can be tracked at scale."

If there's any silver lining for AI companies, it's that model collapse can be slowed down by infusing more original human data into a training set, according to the study. Still, the fact remains: AI models are hungry, and they need high-quality and original data. Can AI companies keep up with that demand?

**More on AI training:** [*AI Companies Running out of Training Data after Burning Through Entire Internet*](<https://futurism.com/the-byte/ai-training-data-shortage>)

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)