---
title: "What Is Synthetic Data? Why AI Trained on AI Is the Next Big Thing (and Problem)"
description: "As AI companies start running out of training data, many are looking into synthetic data — but it remains unclear whether it will work."
date: "2024-04-08"
modified: "2024-04-08"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/what-is-synthetic-data"
categories:
  - "Artificial Intelligence"
tags:
  - "ai training"
  - "anthropic"
  - "synthetic data"
  - "the digest"
---

# What Is Synthetic Data? Why AI Trained on AI Is the Next Big Thing (and Problem)

![As AI companies start running out of training data, many are looking into synthetic data — but it remains unclear whether it will work.](<https://futurism.com/wp-content/uploads/2024/04/synthetic-data-explained.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Short Supply

As AI companies start running out of training data, many are looking into so-called "synthetic data" — but it remains unclear whether such a thing will ever work.

As the [*New York Times* explains](<https://www.nytimes.com/2024/04/06/technology/ai-data-tech-companies.html>), synthetic data is — on its face, at least — a simple solution for the growing scarcity and other issues with AI training data. If AI can grow large on data generated *by* AI, it would not only solve the [training data shortage](<https://theconversation.com/researchers-warn-we-could-run-out-of-data-to-train-ai-by-2026-what-then-216741>), but could also eliminate the looming problem of AI copyright infringement, too.

But while companies like Anthropic, Google, and OpenAI are all working to try to [create quality synthetic data](<https://www.wsj.com/tech/ai/ai-training-data-synthetic-openai-anthropic-9230f8d8>), none have managed to do so quite yet.

Thus far, AI models built on synthetic data [have tended to run into trouble](<https://futurism.com/ai-trained-ai-generated-data>). Australian AI researcher and podcaster Jathan Sadowski referred to the isssues as "Habsburg AI," a reference to the [deeply-inbred Habsburg dynasty](<https://biomedicalodyssey.blogs.hopkinsmedicine.org/2023/08/habsburgs-all-in-the-family/>) and their [ultra-prominent chins](<https://digitalcommons.unl.edu/cgi/viewcontent.cgi?article=1010&context=ureca>) that signaled their family's penchant for intermarriage.

As [Sadowski tweeted last February](<https://twitter.com/jathansadowski/status/1625245803211272194>), this term describes "a system that is so heavily trained on the outputs of other generative AI's that it becomes an inbred mutant, likely with exaggerated, grotesque features" — much like, well, the [Hapsburg jaw](<https://www.smithsonianmag.com/smart-news/distinctive-habsburg-jaw-was-likely-result-royal-familys-inbreeding-180973688/>).

Last summer, [*Futurism* interviewed](<https://futurism.com/ai-trained-ai-generated-data-interview>) another data researcher, Rice University's Richard G. Baraniuk, about his term for this phenomenon: "Model Autophagy Disorder," or "MAD" for short. It took only five generations of AI inbreeding for the model in the Rice research to "blow up," as the professor put it.

## Synthetic Solutions

The big question: can AI companies figure out a way to make synthetic data that doesn't drive their systems nuts?

As the *NYT* explains, OpenAI and Anthropic — which was, notably, founded by former OpenAI employees who wanted to create more ethical AI — are experimenting with a sort of checks-and-balances system. The first model generates the data, and the second checks the data for accuracy.

Thus far, Anthropic has been the most candid about its use of synthetic data, admitting that it uses a "[constitution](<https://www.anthropic.com/news/constitutional-ai-harmlessness-from-ai-feedback>)" or list of guidelines to train out its two-model system and even that Claude 3, the latest version of its LLM, was trained on "[data we generate internally](<https://twitter.com/Justin_Halford_/status/1764677260555034844>)."

While it's a promising concept, the synthetic data research thus far is anything but — and given that researchers [don't *really* know how AI works](<https://futurism.com/the-byte/nobody-knows-how-ai-works>) to begin with, it's difficult to imagine them figuring out synthetic data anytime soon.

**More on AI conundrums:** [*The Person Who Was in Charge of OpenAI's $175 Million Fund Appears to Be Fake*](<https://futurism.com/the-byte/fake-person-openai-fund>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)