---
title: "In Cringe Video, OpenAI CTO Says She Doesn’t Know Where Sora’s Training Data Came From"
description: "Wondering what data OpenAI used to train its buzzy new text-to-video AI? OpenAI CTO Mira Murati seems to be wondering, too."
date: "2024-03-15"
modified: "2024-03-15"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/video-openai-cto-sora-training-data"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai"
  - "OpenAI"
  - "sora"
  - "text-to-video"
---

# In Cringe Video, OpenAI CTO Says She Doesn’t Know Where Sora’s Training Data Came From

![Wondering what data OpenAI used to train its buzzy new text-to-video AI? OpenAI CTO Mira Murati seems to be wondering, too.](<https://futurism.com/wp-content/uploads/2024/03/video-openai-cto-sora-training-data.jpg>)
*\<em\>Image: Wall Street Journal via YouTube / Futurism\</em\>*

Wondering what data OpenAI used to train its buzzy new text-to-video AI? The company's CTO is similarly unsure.

Mira Murati, OpenAI's longtime chief technology officer, [sat down with *The Wall Street Journal's* Joanna Stern this week](<https://www.wsj.com/video/series/joanna-stern-personal-technology/openai-made-me-crazy-videosthen-the-cto-answered-most-of-my-questions/C2188768-D570-4456-8574-9941D4F9D7E2>) to discuss [Sora](<https://futurism.com/the-byte/openai-cto-releasing-sora-this-year>), the company's forthcoming video-generating AI. About halfway through the 10-minute-long interview, Stern straightforwardly asked Murati where the new model's training data was gleaned from. But Murati, in the most cringe-inducing way possible, couldn't find an answer beyond vague corporate language.

"We used publicly available data and licensed data," Murati responded to the resoundingly simple question.

Stern pushed back with more specific source examples: "So, videos on YouTube?"

"I'm actually not sure about that," said Murati, before rebuffing further queries about whether videos shared to Instagram or Facebook were fed into model.

"You know, if they were publicly available — publicly available to use," the CTO answered, "but I'm not sure. I'm not confident about it."

Stern then inquired about OpenAI's [data training partnership](<https://investor.shutterstock.com/news-releases/news-release-details/shutterstock-expands-partnership-openai-signs-new-six-year>) with the stock image company Shutterstock, asking if videos on the partnered platform were sucked into Sora's training material. And this time? Murati decided to shut down the line of questioning altogether.

"I'm just not going to go into detail about the data that was used," Murati continued. "But it was publicly available or licensed data."

So, in sum, Murati can't tell you exactly where the videos gobbled up by Sora first came from. But rest assured, the sourceless data was definitely, one hundred percent publicly available or licensed. Convincing stuff!

It's a bad look all around for OpenAI, which has drawn wide controversy — not to mention [multiple](<https://futurism.com/the-byte/openai-big-legal-trouble>) copyright lawsuits, [including one from *The New York Times*](<https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec2023.pdf>) — for its data-scraping practices. After all, if the company's CTO can't firmly tell you where its buzziest new model's training data was sourced from, it doesn't exactly communicate a [particular amount of care](<https://futurism.com/openai-content-new-york-times-lawsuit>) for the issue from OpenAI's higher-ups.

https://twitter.com/JoannaStern/status/1768306032466428291

After the interview, Murati reportedly confirmed to the *WSJ* that Shutterstock videos were indeed included in Sora's training set. But when you consider the vastness of video content across the web, any clips available to OpenAI through Shutterstock are likely only a small drop in the Sora training data pond.

Online, reactions to the clip were mixed, with many chalking Murati's close-lipped responses up to a possible lack of candidness.

"So when \*the CTO\* of OpenAI is asked if Sora was trained on YouTube videos, she says 'actually I'm not sure' and refuses to discuss all further questions about the training data," former *LA Times* tech columnist Brian Merchant [wrote in an X-formerly-Twitter post](<https://twitter.com/bcmerchant/status/1768160804678070623>). "Either a rather stunning level of ignorance of her own product, or a lie — pretty damning either way!"

"You're the CTO ma'am," [added](<https://twitter.com/thenetrunna/status/1768113426054713458>) another netizen, "you should know."

Others, meanwhile, jumped to Murati's defense, arguing that if you've ever published anything to the internet, you should be perfectly fine with AI companies gobbling it up.

"Why does it matter? That is the question," said one X user. "I find it insane that people make things public to everyone in the world and then complain when someone uses that public thing. If you want to be private, then be private."

That latter argument, though, speaks to the bizarre new reality that internet users have now found themselves in. Historically, when someone told you to be careful of what you post online, the reasoning was something akin to "you might regret that later" — and not "a multibillion-dollar AI company might turn a profit by vacuuming that Facebook video of you and your family, or a goofy YouTube video you made with your friends, into a generative AI model."

Whether Murati was keeping things close to the vest to avoid more [copyright](<https://futurism.com/ai-image-generators-copyrighted-characters>) litigation or simply just didn't know the answer, people have good reason to wonder where AI data — be it "publicly available and licensed" or not — is coming from. And moving forward, vague corporate mumbling probably isn't going to cut it.

**More on OpenAI and its data:** [*OpenAI Says It's Fine to Vacuum Up Everyone's Content and Charge for It Without Paying Them*](<https://futurism.com/openai-content-new-york-times-lawsuit>)

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)