---
title: "If Even 0.001 Percent of an AI’s Training Data Is Misinformation, the Whole Thing Becomes Compromised, Scientists Find"
description: "Researchers found that if a mere 0.001 percent of the training data of a given LLM is \"poisoned,\" the entire thing falls apart."
date: "2025-01-12"
modified: "2025-01-12"
authors:
  - name: "Victor Tangermann"
    job_title: "Senior Editor"
    link: "https://futurism.com/authors/victor"
url: "https://futurism.com/training-data-ai-misinformation-compromised"
categories:
  - "Artificial Intelligence"
tags:
  - "ai chatbots"
  - "large language models"
  - "medical ai"
---

# If Even 0.001 Percent of an AI’s Training Data Is Misinformation, the Whole Thing Becomes Compromised, Scientists Find

![Researchers found that if a mere 0.001 percent of the training data of a given LLM is "poisoned," the entire thing falls apart.](<https://futurism.com/wp-content/uploads/2025/01/training-data-ai-misinformation-compromised.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

It's no secret that large language models (LLMs) like the ones that power popular chatbots like ChatGPT are surprisingly fallible. Even the most advanced ones still have a [nagging tendency](<https://futurism.com/the-byte/facebook-chatbot-mushroom-foraging>) to contort the truth — and with an unnerving degree of confidence.

And when it comes to medical data, those kinds of discrepancies become a whole lot more serious given that lives may be at stake.

Researchers at New York University have found that if a mere 0.001 percent of the training data of a given LLM is "poisoned," or deliberately planted with misinformation, the entire training set becomes likely to propagate errors.

As detailed in a [paper](<https://www.nature.com/articles/s41591-024-03445-1#Sec7>) published in the journal *Nature Medicine*, first [spotted by *Ars Technica*](<https://arstechnica.com/science/2025/01/its-remarkably-easy-to-inject-new-medical-misinformation-into-llms/>), the team also found that despite being error-prone, corrupted LLMs still perform just as well on "open-source benchmarks routinely used to evaluate medical LLMs" as their "corruption-free counterparts."

In other words, there are serious risks involved in making use of biomedical LLMs, which could easily be overlooked using conventional tests.

"In view of current calls for improved data provenance and transparent LLM development," the team writes in its paper, "we hope to raise awareness of emergent risks from LLMs trained indiscriminately on web-scraped data, particularly in healthcare where misinformation can potentially compromise patient safety."

In an experiment, the researchers intentionally injected "AI-generated medical misinformation" into a commonly used LLM training dataset known as "The Pile," which contains "high-quality medical corpora such as PubMed."

The team generated a total of 150,000 medical articles within just 24 hours, and the results were shocking, demonstrating that it's incredibly easy — and even cheap — to effectively poison LLMs.

"Replacing just one million of 100 billion training tokens (0.001 percent) with vaccine misinformation led to a 4.8 percent increase in harmful content, achieved by injecting 2,000 malicious articles (approximately 1,500 pages) that we generated for just US$5.00," the researchers wrote.

Unlike invasive hijacking attacks that can force LLMs to give up confidential information or [even execute](<https://www.ibm.com/think/topics/prompt-injection>) code, data poisoning doesn't require direct access to the model weights, or the numerical values used to define the strength of connections between neurons in an AI.

In other words, attackers only need to "host harmful information online" to undermine the validity of an LLM, according to the researchers.

The research highlights glaring risks involved in the deployment of AI-based tools, especially in a medical setting. And in many ways, the cat is already out of the bag. Case in point, the [*New York Times* reported](<https://futurism.com/neoscope/medical-ai-doctor-lies-records>) last year that an AI-powered communications platform called MyChart, which automatically drafts replies to patients' questions on behalf of doctors, regularly "hallucinates" untrue entries about a given patient's condition.

In short, the fallible nature of LLMs, particularly when it comes to the medical sector, should be a major cause for concern.

"AI developers and healthcare providers must be aware of this vulnerability when developing medical LLMs," the paper reads. "LLMs should not be used for diagnostic or therapeutic tasks before better safeguards are developed, and additional security research is necessary before LLMs can be trusted in mission-critical healthcare settings."

**More on medical AI:** *[Murdered Insurance CEO Had Deployed an AI to Automatically Deny Benefits for Sick People](<https://futurism.com/neoscope/united-healthcare-claims-algorithm-murder>)*

## Author
I've been at Futurism since 2017, where my role has evolved to encompass design, writing, and increasingly editing. I've always been fascinated by space exploration and advanced transportation, which I've leaned into by interviewing luminaries in those fields while closely following the dimensions of policy and regulation that allow next-generation projects to succeed -- or, sometimes, to fail. I'm also keenly interested in the effects of generative AI on society, policies, and democratic institutions, as well as clean energy, physics and biology, and the vagaries of tech leadership. My work for Futurism has been cited by publications including Ars Technica, Gizmodo, PC Magazine, Jalopnik, Fox News, and the New York Post. I spent my childhood living in locations including Manila, the Philippines, and Geneva, Switzerland, attended McGill University, and now live in Toronto, Canada. Before Futurism I worked at AskMen and a small photography studio. In my free time, I'm an avid gardener, foodie, and craft beer lover, as well as a maker of artisanal hot pepper sauces. I have a magnificent dog named Freida.

### Author social links  
[Bluesky](<https://bsky.app/profile/vtanger.bsky.social>)