---
title: "AI Model Used By Hospitals Caught Making Up Details About Patients, Inventing Nonexistent Medications and Sexual Acts"
description: "Nabla, a popular medical AI transcription tool, is powered by OpenAI's Whisper model that researchers have found is horribly unreliable."
date: "2024-10-30"
modified: "2024-10-30"
authors:
  - name: "Frank Landymore"
    job_title: "Contributing Writer"
    link: "https://futurism.com/authors/flandymore"
url: "https://futurism.com/the-byte/whisper-nabla-hospital-ai-details-patients"
categories:
  - "Artificial Intelligence"
tags:
  - "generative ai"
  - "medical ai"
  - "OpenAI"
  - "the digest"
---

# AI Model Used By Hospitals Caught Making Up Details About Patients, Inventing Nonexistent Medications and Sexual Acts

![Nabla, a popular medical AI transcription tool, is powered by OpenAI's Whisper model that researchers have found is horribly unreliable.](<https://futurism.com/wp-content/uploads/2024/10/hospital-ai-making-up-bizarre-details-patients.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Health Scare

In a new [investigation from *The Associated Press*](<https://apnews.com/article/ai-artificial-intelligence-health-business-90020cdf5fa16c79ca2e5b6c4c9bbb14>), dozens of experts have found that Whisper, an AI-powered transcription tool made by [OpenAI](<https://futurism.com/the-byte/openai-whistleblower-copyrighted-data>), is plagued with frequent hallucinations and inaccuracies, with the AI model often inventing completely unrelated text.

What's even more concerning, though, is who's relying on the tech, according to the *AP***:** despite OpenAI warning that its model shouldn't be used in "high-risk domains," over 30,000 medical workers and 40 health systems are using Nabla, a tool built on Whisper, to transcribe and summarize patient interactions — almost certainly with inaccurate results.

In a medical environment, this could have "really grave consequences," Alondra Nelson, a professor at the Institute for Advanced Study, told the *AP*.

"Nobody wants a misdiagnosis," Nelson said. "There should be a higher bar."

## Whisper Campaign

Nabla chief technology officer Martin Raison told the *AP* that the tool was fine-tuned on medical language. Even so, it can't escape the inherent unreliability of its underlying model.

One machine learning engineer who spoke to the AP said he discovered hallucinations in half of the over 100 hours of Whisper transcriptions he looked at. Another who examined 26,000 transcripts said he found hallucinations in almost all of them.

Whisper performed poorly even with well-recorded, short audio samples, according to a recent study cited by the *AP*. Over millions of recordings, there could be tens of thousands of hallucinations, researchers warned.

Another team of researchers revealed just how egregious these errors can be. Whisper would inexplicably add racial commentary, they found, such as making up a person's race without instruction, and also [invent nonexistent medications](<https://x.com/allisonkoe/status/1797663358461817064>). In other cases, the AI would describe [violent and sexual acts](<https://x.com/allisonkoe/status/1797663739174633807>) that had no basis in the original speech. They even found baffling [instances of YouTuber lingo](<https://x.com/allisonkoe/status/1797664087855530176>), such as a "like and subscribe," being dropped into the transcript.

Overall, nearly 40 percent of these errors were harmful or concerning, the team concluded, because they could easily misrepresent what the speaker had actually said.

https://twitter.com/allisonkoe/status/1797663358461817064

## Ground Truth

The scope of the damage could be immense. According to Nabla, its tool has been used to transcribe an estimated seven million medical visits, the paperwork for all of which could now have pernicious inaccuracies somewhere in the mix.

And worryingly, there's no way to verify if the AI transcriptions are accurate, because the tool deletes the original audio recordings "for data safety reasons," according to Raison. Unless the medical workers themselves kept a copy of the recording, any hallucinations will stand as part of the official record.

"You can't catch errors if you take away the ground truth," William Saunders, a research engineer who quit OpenAI in protest, told the *AP*.

Nabla officials said they are aware that Whisper can hallucinate and are addressing the problem, per the *AP*. Being "aware" of the problem, however, seemingly didn't stop the company from pushing experimental and as yet extremely unreliable tech onto the medical industry in the first place.

**More on AI:** *[After Teen's Suicide, Character.AI Is Still Hosting Dozens of Suicide-Themed Chatbots](<https://futurism.com/suicide-chatbots-character-ai>)*

## Author
At Futurism, my work has often centered on bringing a sense of clarity and insight to complex topics ranging from the regulation of emerging technologies to the esoteric ideologies of Silicon Valley executives, while striving not to lose the poetic sense of awe inspired by often-obscure fields like astrophysics and quantum computing. I broke the story of CNET using AI to produce articles that turned out to be riddled with factual errors and plagiarism — a dam-breaking inflection point, as I've reported, that's inspired copycats and endless discourse while beguiling stakeholders ranging from tech giants to purveyors of spam around the web. My work at Futurism has been cited by publications including CBS News, the Los Angeles Times, Vice, Gizmodo, Engadget, the Verge, and Vanity Fair. I grew up in locales ranging from India to China, and now live in the exotic suburbs of Virginia. In my free time, I'm an avid reader of weird sci-fi literature, an aficionado of East Asian cinema, and, regrettably, a relapsed gamer. Allegedly, I’m working on a debut novel, currently untitled.

### Author social links  
[Bluesky](<https://bsky.app/profile/f-w-l.bsky.social>)