---
title: "Authors Suing OpenAI Will Get to See Its Secret Training Data in Heavily Locked Down Room"
description: "Authors suing OpenAI for copyright infringement are going to get unprecedented access to its training data — but that access is very limited. "
date: "2024-09-28"
modified: "2024-09-28"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/authors-suing-openai-training-data-viewing"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "copyright"
  - "OpenAI"
  - "the digest"
  - "training data"
---

# Authors Suing OpenAI Will Get to See Its Secret Training Data in Heavily Locked Down Room

![Authors suing OpenAI for copyright infringement are going to get unprecedented access to its training data — but that access is very limited. ](<https://futurism.com/wp-content/uploads/2024/09/authors-suing-openai-training-data-viewing.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Secure All

Authors suing OpenAI for copyright infringement are going to get unprecedented access to its training data — but only in a performatively locked down room.

As [*The Hollywood Reporter* reveals](<https://www.hollywoodreporter.com/business/business-news/openai-training-data-inspected-authors-copyright-case-1236011291/>), lawyers for authors Sarah Silverman, Ta-Nehisi Coates, and Paul Tremblay announced in a new court filing this week that they have reached an agreement with OpenAI that will allow the writers' representatives to view the company's training data trove.

This is, notably, the first time OpenAI has allowed an outside party to view its training data, and there are some major caveats.

As the report explains, the training data can only be viewed in a "secure" room at OpenAI's San Francisco headquarters on a locked down computer that does not have access to the internet or any shared networks. No outside electronics will be allowed into the room, and although the reps may be allowed to take notes, making copies of any portion of the data is strictly forbidden — a wild request, we should note, since it's all material created by the public in the first place.

Whoever views the datasets, which were not given a size estimate but are [undeniably massive](<https://www.technologyreview.com/2023/04/19/1071789/openais-hunger-for-data-is-coming-back-to-bite-it/>), will be required to provide identification, give their name in a visitor's log, and sign a non-disclosure agreement, the report adds.

## Head Case

These stipulations, which sound more like the protocols for [viewing state secrets](<https://www.washingtonpost.com/national-security/interactive/2023/scif-room-meaning-classified/>) than looking at a bunch of AI training data, are the latest forte in the protracted battle between these authors and OpenAI that may well end up serving as the legal precedent for AI's use of copyrighted material in the future.

The attorneys representing Coates, Silverman, and Tremblay hail from SF's Joseph Saveri Law Firm and are representing them in a [similar case against Meta](<https://www.hollywoodreporter.com/business/business-news/mark-zuckerberg-deposition-sarah-silverman-meta-1236011815/>) arguing the same thing: that the authors' copyrighted work was used without permission or compensation, resulting in ChatGPT spitting out answers that infringe upon their copyrights.

Twice this year, other claims from these lawsuits have been thrown out.

In February, the [majority of the suit against OpenAI was tossed](<https://www.hollywoodreporter.com/business/business-news/sarah-silverman-openai-lawsuit-claims-judge-1235823924/>) as US District Judge Araceli Martinez-Olguin said that the attorneys' claims of negligence, unjust enrichment, and vicarious copyright infringement were without merit.

Later this year, the same judge [threw out the portion of the suit](<https://www.hollywoodreporter.com/business/business-news/sarah-silverman-lawsuit-openai-copyrighted-novels-1235963290/>) that alleged OpenAI engaged in unfair business practices by using the authors' copyrighted works — though as *THR* notes, the direct copyright infringement claims have remained intact throughout.

As of now, it's unclear when these locked down training data viewing sessions will take place, how long they'll have with the data, and how many people will go in at once.

All the same, we'll be watching this one closely — while not quite envying those who will have to dig through all that data to try to find a smoking gun.

**More on OpenAI:** [*OpenAI Execs Mass Quit as Company Removes Control From Non-Profit Board and Hands It to Sam Altman*](<https://futurism.com/openai-execs-quit>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)