---
title: "MIT Deletes Database That Taught AI to Be Racist, Sexist"
description: "MIT has permanently removed a huge and widely used dataset used for training AI after learning it was full of racist and misogynistic slurs."
date: "2020-07-02"
modified: "2020-07-10"
authors:
  - name: "Dan Robitzski"
    link: "https://futurism.com/authors/danrobitzski"
url: "https://futurism.com/the-byte/mit-deletes-database-taught-ai-racist-sexist"
categories:
  - "Artificial Intelligence"
tags:
  - "algorithmic bias"
  - "artificial intelligence"
  - "mit"
  - "the digest"
---

# MIT Deletes Database That Taught AI to Be Racist, Sexist

![MIT permanently deleted a massive dataset used for training AI after learning that it was full of racist and sexist slurs.](<https://futurism.com/wp-content/uploads/2020/07/mit-deletes-database-taught-ai-racist-sexist.jpg>)
*\<em\>Image: Ujesh via Unsplash\</em\>*

## Data Regression

A machine learning algorithm is only as good as the data it's trained on. Unfortunately, a massive and popular training dataset from MIT taught a bunch of algorithms to use racist and misogynistic slurs.

MIT just took down the offending database, 80 Million Tiny Images, for some much-needed sanitation, *[The Register ](<https://www.theregister.com/2020/07/01/mit_dataset_removed/>)*[reports](<https://www.theregister.com/2020/07/01/mit_dataset_removed/>). The dataset has been used to train image-recognition AI since 2008, but had never been probed for racist or offensive content, meaning a major source of [algorithmic bias](<https://futurism.com/the-byte/robot-journalist-accused-racism>) was flying under the radar.

## Square One

AI learns to interpret and identify objects in pictures after poring over thousands of images that were already labeled. In MIT's dataset, thousands of pictures of Black people — and monkeys — were labeled with the N-word. Pictures of women were labeled with misogynistic slurs. After being trained on that data, AI can perpetuate those prejudices [in the real world](<https://futurism.com/the-byte/cops-arrested-innocent-man-facial-recognition>).

"It is clear that we should have manually screened them," MIT computer scientist and electrical engineer Antonio Torralba told *The Register*. "For this, we sincerely apologize. Indeed, we have taken the dataset offline so that the offending images and categories can be removed."

But MIT then clarified that the dataset is gone forever.

## Giving Up

After attempting to screen out the offending images, MIT decided that the task is simply too difficult for humans — there are simply too many pictures to check.

"Therefore, manual inspection, even if feasible, will not guarantee that offensive images can be completely removed," reads [an MIT statement](<https://groups.csail.mit.edu/vision/TinyImages/>).

**READ MORE:** [MIT apologizes, permanently pulls offline huge dataset that taught AI systems to use racist, misogynistic slurs](<https://www.theregister.com/2020/07/01/mit_dataset_removed/>) \[*The Register*\]

**More on AI bias:** *[Robot Journalist Accused of Racism](<https://futurism.com/the-byte/robot-journalist-accused-racism>)*

## Author
Dan Robitzki is a senior reporter for Futurism, where he likes to cover AI, tech ethics, and medicine. He spends his extra time fencing and streaming games from Los Angeles, California.

### Author social links  
[Twitter](<https://x.com/danrobitzski>)