---
title: "Nvidia Caught Stealing Mind-Boggling Quantity of YouTube Videos to Train AI"
description: "AI chip giant Nvidia has been quietly scraping astronomical amounts of YouTube video data to train its AI models, leaked docs reveal."
date: "2024-08-05"
modified: "2024-08-05"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/the-byte/nvidia-caught-scraping-youtube-ai"
categories:
  - "Artificial Intelligence"
tags:
  - "ai"
  - "ai training"
  - "NVIDIA"
  - "the digest"
---

# Nvidia Caught Stealing Mind-Boggling Quantity of YouTube Videos to Train AI

![AI chip giant Nvidia has been quietly scraping astronomical amounts of YouTube video data to train its AI models, leaked docs reveal.](<https://futurism.com/wp-content/uploads/2024/08/nvidia-caught-scraping-youtube-ai.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

Leaked documents [obtained by *404 Media*](<https://www.404media.co/nvidia-ai-scraping-foundational-model-cosmos-project/>) reveal that AI-powering chip giant Nvidia has been quietly scraping astronomical numbers of YouTube video data to train its AI models — a legally and ethically murky decision that adds to the ever-growing pile of deeply questionable, and often very secretive, AI training practices by entities ranging from startups to corporate giants.

According to *404*'s explosive scoop, Nvidia has obtained an eye-watering amount of YouTube data to train AI models including its Cosmos deep learning model, a self-driving car algorithm, a "digital human" AI avatar product, and its 3D world-building tool called Omniverse.

Nvidia also reportedly took pains to hide its activities from YouTube, using dozens of "virtual machines" that automatically changed their IP addresses to avoid detection.

Neither individual video creators nor YouTube owner Google, a [notable Nvidia customer](<https://nvidianews.nvidia.com/news/google-cloud-ai-development>), consented to Nvidia's data scraping. And internal correspondence between Nvidia employees, including from its higher-ups, reveals a wildly brash, ask-questions-later — or ask-questions-hopefully-never — approach to the covert data-vacuuming campaign.

"We are finalizing the v1 data pipeline and securing the necessary computing resources," Ming-Yu Liu, Nvidia's VP of Research and a leader on the Cosmos project, wrote in a May email, according to *404*, "to build a video data factory that can yield a human lifetime visual experience worth of training data per day."

What's more, in response to employee concerns regarding the legality and ethics of Nvidia's newfound data acquisition practices, managers including Liu insisted that the move was approved from the top down.

"This is an executive decision," Liu wrote to a hesitant underling on one such occasion, according to Slack messages reviewed by *404*. "We have an umbrella approval for all of the data."

In one particularly egregious case, documents obtained by *404* revealed that Nvidia at one point knowingly trained its models on HD-VG-130M, a dataset trained on 130 million YouTube videos created explicitly for academic research. Given that Nvidia was using that academic data to train commercial models, its a horrible look.

"I think there's a huge gap between commercializing something without someone's consent," Shayne Longpre, a PhD Candidate at the MIT Media Lab, told *404* of the misuse of research-intended data, "versus studying the generative AI capabilities based off of things that have been publicly put online."

Nvidia has emerged as a [central player in the AI industry](<https://www.reuters.com/technology/artificial-intelligence/delay-nvidias-new-ai-chip-could-affect-microsoft-google-meta-information-says-2024-08-03/>) due to its market dominance over graphic processing units (GPUs), which are the computing chips that often support compute-heavy AI systems. AI companies including OpenAI, Microsoft, Meta, and — again — Google count themselves as Nvidia customers, rendering Nvidia's sneaky use of what ultimately is Google-owned data all the more scandalous. Every major player in the AI industry is battling it out for dominance — including Nvidia, the market's hardware backbone, and now a proven frenemy.

Indeed, when asked by *404* about Nvidia's scraping practices, a spokesperson for Google pointed to an April interview in which YouTube CEO Neal Mohan [told *Bloomberg*](<https://www.bloomberg.com/news/articles/2024-04-04/youtube-says-openai-training-sora-with-its-videos-would-break-the-rules?embedded-checkout=true&ref=404media.co>) that [using YouTube's data](<https://futurism.com/the-byte/openai-executive-choked-sora-youtube>) without permission is in "clear violation" of the platform's terms of service.

"When a creator uploads their hard work to our platform, they have certain expectations," Mohan told *Bloomberg*. "One of those expectations is that the terms of service is going to be abided by. It does not allow for things like transcripts or video bits to be downloaded, and that is a clear violation of our terms of service."

In a statement to *404*, Nvidia claimed that its AI training practices are "in full compliance with the letter and the spirit of copyright law." The jury's still out, of course, on how the humans who made the allegedly lifetimes' worth of content now powering the chip maker's AI systems feel about that.

**More on Nvidia:** [*Is the Tech Stock Collapse Related a Sign of the AI Bubble Popping?*](<https://futurism.com/the-byte/tech-stocks-ai>)

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)