---
title: "AI Appears to Rapidly Be Approaching Brick Wall Where It Can’t Get Smarter"
description: "Researchers warn that companies like OpenAI and Google could soon run out of human-written training data for their AI models."
date: "2024-06-08"
modified: "2024-06-08"
authors:
  - name: "Victor Tangermann"
    job_title: "Senior Editor"
    link: "https://futurism.com/authors/victor"
url: "https://futurism.com/the-byte/ai-running-out-data-smarter"
categories:
  - "Artificial Intelligence"
tags:
  - "anthropic"
  - "artificial intelligence"
  - "OpenAI"
  - "the digest"
---

# AI Appears to Rapidly Be Approaching Brick Wall Where It Can’t Get Smarter

![Researchers warn that companies like OpenAI and Google could soon run out of human-written training data for their AI models.](<https://futurism.com/wp-content/uploads/2024/06/ai-running-out-data-smarter.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Devourer

Researchers are ringing the alarm bells, warning that companies like OpenAI and Google are rapidly running out of human-written training data for their AI models.

And without new training data, it's likely the models won't be able to get any smarter, a point of reckoning for the burgeoning AI industry.

"There is a serious bottleneck here," AI researcher Tamay Besiroglu, lead author of a [new paper](<https://arxiv.org/abs/2211.04325>) to be presented at a conference this summer, [told the *Associated Press*](<https://apnews.com/article/ai-artificial-intelligence-training-data-running-out-9676145bac0d30ecce1513c20561b87d>). "If you start hitting those constraints about how much data you have, then you can’t really scale up your models efficiently anymore."

"And scaling up models has been probably the most important way of expanding their capabilities and improving the quality of their output," he added.

## Feed Me

It's an existential threat for AI tools that rely on feasting on copious amounts of data, which has often indiscriminately been pulled from publicly available archives online.

The controversial trend has already led to publishers, [including the *New York Times*](<https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html>), suing OpenAI over copyright infringement for using their material to train AI models.

And as companies continue to [lay off workers while making major investments in AI](<https://futurism.com/the-byte/microsoft-layoffs-blaming-ai-wave>), the stream of new content could soon turn into a trickle.

The latest paper, authored by researchers at San Francisco-based think tank Epoch, suggests that the sheer amount of text data AI models are being trained on is growing roughly 2.5 times a year. Meanwhile, computing has outpaced that considerably, growing four times a year.

Extrapolated on a graph, that means large language models like Meta's Llama 3 or OpenAI's GPT-4 could entirely run out of fresh data as soon as 2026, the researchers argue.

## AI Ouroboros

Once AI companies do run out of training data, something that's [been predicted by other researchers as well](<https://futurism.com/the-byte/ai-training-data-shortage>), they're likely to try training their large language models on AI-generated data instead. Outfits including OpenAI, Google, and Anthropic are already [working on ways to generate "synthetic data](<https://futurism.com/the-byte/what-is-synthetic-data>)" for this purpose.

It's not clear that will work, according to experts. In a [paper](<https://arxiv.org/pdf/2307.01850>) last year, scientists at Rice and Stanford University found that feeding their models AI-generated content [causes their output quality to erode](<https://futurism.com/ai-trained-ai-generated-data>). Both language models and image generators could be sent down an "autophagous loop," the AI equivalent of a snake eating its own tail.

But whether any of this will actually become a problem remains a subject of debate. For one, we'd be perfectly fine *without* [wasting copious amounts of energy](<https://futurism.com/the-byte/ai-electricity-use-spiking-power-entire-country>) and water on training these AIs.

And it's possible that AI algorithms themselves will become more efficient, producing better outputs with less training data or computing power.

"I think it’s important to keep in mind that we don’t necessarily need to train larger and larger models," AI researcher and University of Toronto assistant professor of computer engineering Nicolas Papernot, who was not involved in the study, told the *AP*.

**More on AI:** *[Sam Altman Admits That OpenAI Doesn't Actually Understand How Its AI Works](<https://futurism.com/sam-altman-admits-openai-understand-ai>)*

## Author
I've been at Futurism since 2017, where my role has evolved to encompass design, writing, and increasingly editing. I've always been fascinated by space exploration and advanced transportation, which I've leaned into by interviewing luminaries in those fields while closely following the dimensions of policy and regulation that allow next-generation projects to succeed -- or, sometimes, to fail. I'm also keenly interested in the effects of generative AI on society, policies, and democratic institutions, as well as clean energy, physics and biology, and the vagaries of tech leadership. My work for Futurism has been cited by publications including Ars Technica, Gizmodo, PC Magazine, Jalopnik, Fox News, and the New York Post. I spent my childhood living in locations including Manila, the Philippines, and Geneva, Switzerland, attended McGill University, and now live in Toronto, Canada. Before Futurism I worked at AskMen and a small photography studio. In my free time, I'm an avid gardener, foodie, and craft beer lover, as well as a maker of artisanal hot pepper sauces. I have a magnificent dog named Freida.

### Author social links  
[Bluesky](<https://bsky.app/profile/vtanger.bsky.social>)