---
title: "No One Wants Apple To Scrape Their Websites for AI Training"
description: "A slew of major news publishers and top social media websites are blocking Apple from scraping their websites for AI training purposes."
date: "2024-09-01"
modified: "2024-09-01"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/the-byte/apple-ai-training"
categories:
  - "Artificial Intelligence"
tags:
  - "ai"
  - "ai training"
  - "apple ai"
  - "the digest"
---

# No One Wants Apple To Scrape Their Websites for AI Training

![A slew of major news publishers and top social media websites are blocking Apple from scraping their websites for AI training purposes.](<https://futurism.com/wp-content/uploads/2024/08/apple-ai-training.jpg>)
*\<em\>Image: Ying Tang / NurPhoto via Getty / Futurism\</em\>*

## Stop Sign

*Wired* [reports that](<https://www.wired.com/story/applebot-extended-apple-ai-scraping/>) a slew of major websites, including influential news publishers and top social media platforms, are blocking Apple's web crawler from scraping their pages for AI training content.

Per the report, media companies that have altered their robots.txt files to lock Applebot out include *The New York Times*, *The Atlantic*, *The Financial Times*, Gannett, Vox Media, and Condé Nast. On the social media side, Facebook, Instagram, and Tumblr all confirmed that they've blocked Apple from scraping their sites, as did the enduring internet elder Craiglist.

Robots.txt files are becoming an increasingly fascinating place to study the digital politics of AI. Some of these companies — including Vox, Condé Nast, and *The Atlantic* — have inked content licensing deals with OpenAI; *The New York Times*, meanwhile, [has drawn a clear line in the sand on AI](<https://futurism.com/chatgpt-plagiarized-nyt-articles>), and is actively [suing](<https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html>) OpenAI for copyright infringement. Facebook and Instagram are both owned by Meta, one of Apple's competitors in the AI field, while platforms built on user content like Tumblr and Craigslist are sitting on some very lucrative troves of quality data. Meanwhile, in the background, Apple has already [entered a deal](<https://openai.com/index/openai-and-apple-announce-partnership/>) *with* OpenAI to integrate the chatbot ChatGPT into "Apple experiences."

In short, the AI industry is intensely competitive, particularly regarding [access](<https://futurism.com/ai-companies-training-data>) to high-quality, [human-made](<https://futurism.com/the-byte/ai-trained-with-ai-generated-data-gibberish>) training material. And as the tentative bonds between AI companies and data wells like journalistic bodies or social media sites continue to take shape, where and how bots like Apple's are allowed to roam offers an interesting glimpse into AI-focused decision-making — on the publisher side, and on behalf of AI companies as well.

## Coal Mines

According to *Wired*, these websites have specifically blocked "Apple-Extended," a web crawler that, [per an Apple blog post](<https://support.apple.com/en-us/119829>), explicitly provides web publishers with the choice to "opt out of their website content being used to train Apple’s foundation models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer Tools." An Apple spokesperson confirmed to *Wired* that blocking Applebot-Extended doesn't ward off the OG Applebot from trawling a website, and instead prevents any scraped data from being used to train Apple's AI models.

Applebot, in contrast, scrapes data for Apple's Siri and Spotlight — a distinction that seems to speak to some caution on Apple's behalf regarding copyright and IP protection in the AI era.

The *NYT* isn't the only company or group [suing AI makers](<https://news.bloomberglaw.com/ip-law/the-intercept-bolsters-openai-copyright-suit-with-more-evidence>), and it may well be in Apple's best interest to avoid scraping any controversial or currently-in-litigation data, especially if it's already tapped OpenAI to fill in some of its product gaps. Call it the billion-dollar canary in the coal mine.

**More on AI and copyright:** [*Amid New York Times Lawsuit, ChatGPT Is Citing Plagiarized Versions of NYT Articles on an Armenian Content Mill*](<https://futurism.com/chatgpt-plagiarized-nyt-articles>)

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)