---
title: "Microsoft’s AI Team Accidentally Leaks Terabytes of Company Data"
description: "Microsoft AI researchers accidentally leaked 38 terabytes of confidential company data — all because of one misconfigured permission token."
date: "2023-09-19"
modified: "2023-09-19"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/the-byte/microsofts-ai-leaks-company-data"
categories:
  - "Artificial Intelligence"
tags:
  - "ai"
  - "data leak"
  - "microsoft"
  - "the digest"
---

# Microsoft’s AI Team Accidentally Leaks Terabytes of Company Data

![Microsoft AI researchers accidentally leaked 38 terabytes of confidential company data — all because of one misconfigured permission token.](<https://futurism.com/wp-content/uploads/2023/09/microsofts-ai-leaks-company-data.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Uh Oh

"Oops" doesn't even cover this one.

Microsoft AI researchers accidentally leaked a staggering 38 terabytes — yes, *terabytes —* of confidential company data on the developer site GitHub, a [new report](<https://www.wiz.io/blog/38-terabytes-of-private-data-accidentally-exposed-by-microsoft-ai-researchers>) from cloud security company Wiz has revealed.

The scope of the data spill is extensive, to say the least. Per the report, the leaked files contained a full disc backup of two employees' workstations, which included sensitive personal data along with company "secrets, private keys, passwords, and over 30,000 internal Microsoft Teams messages."

Worse yet, the leak could have even made Microsoft's AI systems vulnerable to cyberattacks.

In short, it's a huge mess — and somehow, it all goes back to one misconfigured URL, a reminder that human error can have some devastating consequences, particularly in the burgeoning world of AI tech.

https://twitter.com/hillai/status/1703771673411871227?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E1703771673411871227%7Ctwgr%5E448c466bba035b8093f1a575f7827072e6f56057%7Ctwcon%5Es1_&ref_url=https%3A%2F%2Fmashable.com%2Farticle%2Fmicrosoft-ai-researchers-leaked-private-data-azure-link-github

## Treasure Trove

According to Wiz, the mistake was made when Microsoft AI researchers were attempting to publish a "bucket of open-source training material" and "AI models for image recognition" to the developer platform.

The researchers miswrote the files' accompanying SAS token, or the storage URL that establishes file permissions. Basically, instead of granting GitHub users access to the downloadable AI material specifically, the butchered token allowed general access to the entire storage account.

And we're not just talking read-only permissions. The mistake actually granted "full control" access, meaning that anyone who might have wanted to tinker with the many terabytes of data — including that of the AI training material and AI models included in the pile — would have been able to.

An "attacker could have injected malicious code into all the AI models in this storage account," Wiz's researchers write, "and every user who trusts Microsoft’s GitHub repository would've been infected by it."

The Wiz report also notes that the SAS misconfiguration *dates back to* *2020*, meaning that this sensitive material has basically been open-season for several years.

## Bad Week

Microsoft says that it's since resolved the issue, writing in a [Monday blog post](<https://msrc.microsoft.com/blog/2023/09/microsoft-mitigated-exposure-of-internal-information-in-a-storage-account-due-to-overly-permissive-sas-token/>) that no customer data was exposed in the leak.

Regardless, this is shaping up to be a terrible week for the Silicon Valley giant, as [reports revealed this morning](<https://www.nbcnews.com/tech/video-games/microsofts-xbox-plans-revealed-emails-tied-ftc-case-rcna105766>) that yet another Microsoft leak — this one related to the company's ongoing battle with the FTC over its attempted acquisition of Activision Blizzard — exposed the company's plans for its next-generation Xbox, in addition to a slew of confidential company correspondence and information.

If there's any takeaway, according to Wiz, it's simply that handling the massive amounts of data required to train AI models demand high levels of care and security precautions, especially as companies rush new AI products to market.

**More on consequential mistakes:** [*Casinos Shut down amid Hacker Intrusions*](<https://futurism.com/the-byte/casinos-shuts-down-hackers>)

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)