---
title: "OpenAI Whistleblower Disgusted That His Job Was to Vacuum Up Copyrighted Data to Train Its Models"
description: "A former OpenAI staffer is blowing the whistle on the company's AI training practices, alleging they violated copyright law."
date: "2024-10-24"
modified: "2024-10-24"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/the-byte/openai-whistleblower-copyrighted-data"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai"
  - "ai copyright"
  - "OpenAI"
  - "the digest"
---

# OpenAI Whistleblower Disgusted That His Job Was to Vacuum Up Copyrighted Data to Train Its Models

![A former OpenAI staffer is blowing the whistle on the company's AI training practices, alleging they violated copyright law.](<https://futurism.com/wp-content/uploads/2024/10/openai-whistleblower-copyrighted-data.jpg>)
*TURIN, ITALY - SEPTEMBER 25: Sam Altman Co-founder and CEO of OpenAI speaks during the Italian Tech Week 2024 at OGR Officine Grandi Riparazioni on September 25, 2024 in Turin, Italy. (Photo by Stefano Guidi/Getty Images) \<em\>Image: Stefano Guidi/Getty Images\</em\>*

## Sounding the Alarm

A former OpenAI researcher is blowing the whistle on the company's AI training practices, alleging that OpenAI violated copyright law to train its AI models — and arguing that OpenAI's current business model stands to upend the business of the internet as we know it, [according to ](<https://www.nytimes.com/2024/10/23/technology/openai-copyright-law.html>)*[The New York Times](<https://www.nytimes.com/2024/10/23/technology/openai-copyright-law.html>).*

The ex-staffer, a 25-year-old named Suchir Balaji, worked at OpenAI for four years before deciding to leave the AI firm due to ethical concerns. As Balaji sees it, because ChatGPT and other OpenAI products have become so heavily commercialized, OpenAI's practice of scraping online material en masse to feed its data-hungry AI models no longer satisfies the criteria of the fair use doctrine. OpenAI — which is currently facing several copyright lawsuits, including a [high-profile case](<https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html>) brought last year by the *NYT* — has [argued the opposite](<https://futurism.com/openai-content-new-york-times-lawsuit>).

"If you believe what I believe," Balaji told the *NYT*, "you have to just leave the company."

Balaji's warnings, which he [outlined in a post](<https://suchir.net/fair_use.html>) on his personal website yesterday, add to the ever-growing controversy around the AI industry's collection and use of copyrighted material to [train AI models](<https://futurism.com/the-byte/zuckerberg-fine-train-ai-data-no-value>), which was largely conducted [without comprehensive government regulation](<https://www.weforum.org/agenda/2024/05/why-regulating-ai-can-be-surprisingly-straightforward-providing-you-have-eternal-vigilance/>) and outside of the public eye.

"Given that AI is evolving so quickly," intellectual property lawyer Bradley Hulbert told the *NYT*, "it is time for Congress to step in."

## Flipping the Switch

Balaji, who was hired in 2020, was one of several staffers tasked with collecting and organizing web-gathered training data that would eventually be fed into OpenAI's large language models (LLMs). Because OpenAI was still technically just a well-funded *research* company at the time, the issue of copyright wasn't [as big of a deal](<https://futurism.com/the-byte/nvidia-caught-scraping-youtube-ai>).

"With a research project, you can, generally speaking, train on any data," Balaji told the *NYT*. "That was the mindset at the time."

But once ChatGPT was released in November 2022, Balaji says, his feelings started to change. After all, the chatbot was no longer a closed-door research project; instead, powered by OpenAI's LLMs, it was being commodified for commercial use — including in cases where the AI was being used to produce content or services that directly reflected or mimicked the copyrighted source material it was trained on, thus [threatening](<https://futurism.com/former-google-news-director-big-tech-journalism>) the livelihoods and profit models of those very individuals and businesses.

"This is not a sustainable model," Bilaji told the *NYT*, "for the internet ecosystem as a whole."

For its part, in a statement to the *NYT*, OpenAI — which has [since abandoned](<https://futurism.com/the-byte/openai-most-orwellian-company>) its non-profit roots [entirely](<https://futurism.com/openai-sleazy-company-creating-agi>) — argued that it builds its "AI models using publicly available data, in a manner protected by fair use and related principles" **and** that is "critical for "US competitiveness."

**More on OpenAI:** [*OpenAI Pivoting From "Benefiting Humanity" to "Making Lots of Money"*](<https://futurism.com/openai-pivoting-benefiting-humankind-making-money>)

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)