---
title: "AI Companies Running Out of Training Data After Burning Through Entire Internet"
description: "AI companies are swiftly running into a massive problem: there isn't enough data on the internet to train the next generation of models."
date: "2024-04-01"
modified: "2024-04-01"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/ai-training-data-shortage"
categories:
  - "Artificial Intelligence"
tags:
  - "ai training"
  - "anthropic"
  - "OpenAI"
  - "the digest"
---

# AI Companies Running Out of Training Data After Burning Through Entire Internet

![AI companies are swiftly running into a massive problem: there isn't enough data on the internet to train the next generation of models.](<https://futurism.com/wp-content/uploads/2024/04/ai-training-data-shortage.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Mass Shortage

As AI companies keep building [bigger and better models](<https://www.quantamagazine.org/the-unpredictable-abilities-emerging-from-large-ai-models-20230316/>), they're running down a shared problem: sometime soon, the internet won't be big enough to provide all the data they need.

As the [*Wall Street Journal* reports](<https://www.wsj.com/tech/ai/ai-training-data-synthetic-openai-anthropic-9230f8d8>), some companies are looking for alternative sources of data training now that the internet is growing too small, with things like publicly-available video transcripts and even AI-generated "synthetic data" as options.

While there are some companies, such as Dataology, which was formed by ex-Meta and Google DeepMind researcher Ari Morcos, looking into ways to train larger and smarter models with less data and resources, most big companies are looking into novel — and controversial — means of data training.

OpenAI, for instance, has per the *WSJ*'s sources discussed training GPT-5 on transcriptions from public YouTube videos — even as its own chief technology officer, Mira Murati, [struggles to answer questions](<https://futurism.com/video-openai-cto-sora-training-data>) about whether its Sora video generator was trained using YouTube data.

## Don't Panic

Synthetic data, meanwhile, has been the subject of ample debate in recent months after researchers found last year that training an AI model on AI-generated data would be a digital form of "[inbreeding](<https://twitter.com/jathansadowski/status/1625245803211272194>)" that would ultimately lead to "[model collapse](<https://futurism.com/ai-trained-ai-generated-data>)" or "[Habsburg AI](<https://futurism.com/ai-trained-ai-generated-data-interview>)."

Some companies, like OpenAI and Anthropic, which was formed by OpenAI in 2021 in efforts to [build a safer and more ethical AI](<https://www.vox.com/future-perfect/23794855/anthropic-ai-openai-claude-2>) than those of their former employer, are seeking to head that off by creating supposedly higher-quality synthetic data — though of course, neither is letting press in on the secret sauce of what exactly that would entail.

Indeed, Anthropic admitted when [announcing its Claude 3 LLM](<https://twitter.com/Justin_Halford_/status/1764677260555034844>) that the model was trained on "data we generate internally," and in an interview with *WSJ*, chief company scientist Jared Kaplan said that he thinks there are good use cases for synthetic data as well.

While concerns about AI running out of data seem to have been [spooking researchers](<https://hbr.org/2019/01/the-future-of-ai-will-be-about-less-data-not-more>) for [some time](<https://futurism.com/ai-companies-training-data>), researcher Pablo Villalobos told the newspaper that although his firm, Epoch, has estimated that AI will run out of usable training data within the next few years, there's no reason for panic.

"The biggest uncertainty," Villalobos said, "is what breakthroughs you’ll see."

Then again, there is another obvious solution to this manufactured problem: AI companies could simply stop trying to create bigger and better models, given that aside from the training data shortage, they also use [tons of electricity](<https://www.theverge.com/24066646/ai-electricity-energy-watts-generative-consumption>) and expensive computing chips that require the [mining of rare-earth minerals](<https://theconversation.com/demand-for-computer-chips-fuelled-by-ai-could-reshape-global-politics-and-security-224438>).

**More on AI training:** [*Microsoft and OpenAI Reportedly Building $100 Billion Secret Supercomputer to Train Advanced AI*](<https://futurism.com/the-byte/microsoft-openai-supercomputer>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)