---
title: "YouTubers Furious After Apple and Anthropic Steal Their Data to Train AI"
description: "A giant dataset of YouTube subtitles has been used to train AI without the permission of the creators whose work was scraped."
date: "2024-07-17"
modified: "2024-07-17"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/youtubers-apple-anthropic-data-ai"
categories:
  - "Anthropic"
  - "Artificial Intelligence"
tags:
  - "ai training"
  - "apple"
  - "the digest"
  - "YouTube"
---

# YouTubers Furious After Apple and Anthropic Steal Their Data to Train AI

![A giant dataset of YouTube subtitles has been used to train AI without the permission of the creators whose work was scraped.](<https://futurism.com/wp-content/uploads/2024/07/youtubers-apple-anthropic-data-ai.jpg>)
*\<em\>Image: Nic Coury / AFP via Getty\</em\>*

## Dirty Work

A giant dataset of YouTube subtitles has, per a new investigation, been used to train countless AI models without the permission of the tens of thousands of creators whose work was scraped.

As [*Wired* reports](<https://www.wired.com/story/youtube-training-data-apple-nvidia-anthropic/>) with the help of the [data-driven *Proof News* project](<https://www.proofnews.org/apple-nvidia-anthropic-used-thousands-of-swiped-youtube-videos-to-train-ai/>), a dataset known as "YouTube Subtitles" has been used by everyone from Apple and Anthropic to Nvidia and Salesforce to train AI models since it was released in 2020.

Compiled by the open-source nonprofit EleutherAI, the YouTube Subtitles dataset doesn't include any actual video, but instead subtitle data from 173,536 videos gleaned from more than 48,000 channels. Among those channels were everything from MIT and Harvard to MrBeast and the *BBC*, among many others.

Of all the channel owners that *Proof* managed to speak with for the story, none had been made aware ahead of time that ElutherAI had used subtitles from their videos.

## Forgiveness, Not Permission

One of the impacted creators, the progressive vlogger David Pakman, was mighty peeved when he learned from *Proof* about his videos being included in the dataset.

"No one came to me and said, 'We would like to use this,'" the commentator, who had nearly 16o videos used in the dataset, told *Wired*. "This is my livelihood, and I put time, resources, money, and staff time into creating this content."

According to AI policy researcher Jai Vipra of Brazil's Fundação Getulio Vargas Law School, the YouTube Subtitles dataset is a "gold mine" because it can teach models how to replicate human speech.

To science vlogger Dave Farina of the popular "Professor Dave Explains" series, however, that gold mine comes at a cost to creators.

"It's still the sheer principle of it," Farina told *Wired*. "If you’re profiting off of work that I’ve done that will put me out of work or people like me out of work, then there needs to be a conversation on the table about compensation or some kind of regulation."

When *Proof* reached out to YouTube owner Google, EleutherAI, and the companies that had used the dataset, only a Google spokesperson chose to respond publicly to say that the company has taken "action over the years to prevent abusive, unauthorized scraping."

It's a provocative state of affairs — and it's hard to tell at this juncture how to fix it if companies won't even speak on the record about it.

**More on AI data:** [*AI Is Being Trained on Images of Real Kids Without Consent*](<https://futurism.com/ai-trained-images-kids>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)