---
title: "OpenAI Sued for Using Everybody’s Writing to Train AI"
description: "A new lawsuit is alleging that OpenAI violated individual data privacy and copyright by using scraped data to train its AI language model."
date: "2023-06-29"
modified: "2023-06-29"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/the-byte/openai-sued-train-ai"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai"
  - "data privacy"
  - "OpenAI"
  - "the digest"
---

# OpenAI Sued for Using Everybody’s Writing to Train AI

![A new lawsuit is alleging that OpenAI violated individual data privacy and copyright by using scraped data to train its AI language model.](<https://futurism.com/wp-content/uploads/2023/06/openai-sued-train-ai.jpg>)
*WASHINGTON, DC - MAY 16: Samuel Altman, CEO of OpenAI, appears for testimony before the Senate Judiciary Subcommittee on Privacy, Technology, and the Law May 16, 2023 in Washington, DC. The committee held an oversight hearing to examine A.I., focusing on rules for artificial intelligence. (Photo by Win McNamee/Getty Images) \<em\>Image: Win McNamee/Getty Images\</em\>*

## My Data, Your Data

A new [lawsuit](<https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rIZH4FXwShJE/v0>) against ChatGPT creator OpenAI is alleging that the buzzy Silicon Valley firm's AI training practices violated the privacy and copyright of — well, of pretty much everyone who's [ever posted anything online](<https://futurism.com/ai-fake-quotes-real-people>).

To train its powerful AI language models, OpenAI utilized an incredible amount of data scraped from various corners of the web. Although [OpenAI doesn't even know](<https://www.newyorker.com/news/daily-comment/what-we-still-dont-know-about-how-ai-is-trained>) exactly what its systems are trained on, those datasets include everything from Wikipedia articles and famous novels to social media posts and [incredibly niche erotica](<https://futurism.com/chat-gpt-sex-omegaverse>) — and OpenAI [didn't ask permission](<https://www.theverge.com/2023/5/5/23709833/openai-chatgpt-gdpr-ai-regulation-europe-eu-italy>) for any of it.

The class action suit, filed in California, alleges that failing to follow proper procurement guidelines, including seeking the consent of those who produced that content in the first place, amounts to straight-up data theft.

"Despite established protocols for the purchase and use of personal information, Defendants took a different approach: theft," reads the filing. "They systematically scraped 300 billion words from the internet, 'books, articles, websites and posts — including personal information obtained without consent.'"

## Not So Free Web

It's a fair criticism. If you've been online at all in the past few decades, your digital outputs are likely embedded into OpenAI's datasets, meaning that anything that OpenAI's generative models churn out — [for profit](<https://futurism.com/the-byte/openai-billions-bad-ai>) — might have bits and pieces of your silently-scraped digital labor embedded into it.

"All of that information is being taken at scale," Ryan Clarkson, the managing partner at the firm suing OpenAI, [told *The Washington Post*](<https://www.washingtonpost.com/technology/2023/06/28/openai-chatgpt-lawsuit-class-action/>), "when it was never intended to be utilized by a large language model."

That said, whether the case actually holds up in court remains to be seen. The internet's infrastructure is complicated, and what's largely seen as the free and open web is often neither of those things; platforms have their own user terms and agreements, and even if we're the ones who have done the work to pack those sites with content, in many cases it [technically belongs to the platform](<https://fortune.com/2022/01/28/big-tech-data-privacy-ethicaltech/>) — and not, unfortunately, to the users.

"When you put content on a social media site or any site, you're generally granting a very broad license to the site to be able to use your content in any way," Katherine Gardner, an intellectual-property lawyer, told *WaPo*. "It's going to be very difficult for the ordinary end user to claim that they are entitled to any sort of payment or compensation for use of their data as part of the training."

**More on OpenAI lawsuits:** [*Man Sues OpenAI after ChatGPT Claimed He Embezzled Money*](<https://futurism.com/the-byte/man-sues-openai-chatgpt-claimed-embezzled-money>)

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)