---
title: "ChatGPT Can Pass Medical Tests, But Its Actual Medical Advice Is a Lot More Dubious"
description: "When Stanford researchers compared ChatGPT's medical advice on real scenarios compared to what medical experts said, they discovered a notable discrepancy."
date: "2023-04-03"
modified: "2023-04-03"
authors:
  - name: "Frank Landymore"
    job_title: "Contributing Writer"
    link: "https://futurism.com/authors/flandymore"
url: "https://futurism.com/neoscope/chatgpt-medical-advice-dubious"
categories:
  - "Developments"
  - "Health & Medicine"
  - "Medical"
tags:
  - "chatgpt"
  - "large language models"
  - "the digest"
---

# ChatGPT Can Pass Medical Tests, But Its Actual Medical Advice Is a Lot More Dubious

![When Stanford researchers compared ChatGPT's medical advice on real scenarios compared to what medical experts said, they discovered a notable discrepancy.](<https://futurism.com/wp-content/uploads/2023/04/chatgpt-medical-advice-dubious.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

A few months ago, OpenAI CEO Sam Altman posited that [AIs like ChatGPT could serve as a "medical advisor"](<https://futurism.com/the-byte/openai-ceo-ai-medical-advice>) for poor people without healthcare.

It sounded like a dumb idea then, and it's sounding like a dumb idea now. In fact, according to [new research](<https://hai.stanford.edu/news/how-well-do-large-language-models-support-clinician-information-needs>) from medical experts at Stanford University, even though OpenAI's ChatGPT has passed [all sorts of tests](<https://futurism.com/the-byte/gpt-4-exam-scores>) including [the US Medical Licensing Exam](<https://healthitanalytics.com/news/chatgpt-passes-us-medical-licensing-exam-without-clinician-input>), the chatbot is worryingly unreliable at responding to real life medical scenarios.

*STAT News* [reports that](<https://www.statnews.com/2023/04/03/gpt-4-chatgpt-health-care-medical-exams/>) the research, while not yet fully released and still pending peer review, found that nearly 60 percent of ChatGPT's answers to actual medical situations either disagreed with a human expert's opinion or weren't relevant enough to be helpful. Not exactly an A+.

In their testing, the Stanford researchers asked the chatbot 64 real life medical questions, recorded its responses, and had twelve clinical experts evaluate them.

With GPT-4, the latest and most powerful large language model that powers the chatbot, over 90 percent of its answers were deemed "safe" enough to not be harmful (though not necessarily completely accurate). Could be worse!

Still, only a considerably lower 41 percent of its answers "agreed" with the answers of medical experts, while some 29 percent were simply too vague or irrelevant to be assessed — which would seem too unreliable for a potential medical assistant, at least for now.

Some have walked back claims of AI's usefulness in this regard, framing it instead as a helpful tool to handle tedious medical paperwork or to provide patients with instructions. But Mark Sendak, a clinical data scientist at Duke University, says this is a slippery slope.

"We shouldn't feel reassured by claims that these tools are only intended to help physicians \[with administrative tasks\]," Sendak told *STAT*, adding that he's doubtful that proper evaluations on AI's ability to handle medical "back of house" tasks will be done consistently.

To be fair, the humans involved in the testing had an advantage: access to patients' health records, which ChatGPT is obviously not privy to. But this in turn highlights an inherent flaw of tests done on the AI, the researchers say. That is, only evaluating it on its textbook knowledge and not its ability to actually help doctors, echoing Sendak's doubts that these AIs will be tested properly.

"We're evaluating these technologies the wrong way," Nigam Shah, a professor of medicine at Stanford who led the research, told *STAT*. "What we should be asking and evaluating is the hybrid construct of the human plus this technology."

Shah adds that he was still "blown away" by GPT-4's improvement over its predecessor, which only agreed with medical experts 20 percent of the time.

Next, he envisions testing ChatGPT's ability to help consult on a tumor board, comparing the results of a board using the AI to one that's entirely human.

**More on ChatGPT:** *[Scammers Are Using ChatGPT to Write Emails That Aren't Riddled With Typos](<https://futurism.com/tags/chatgpt>)*

## Author
At Futurism, my work has often centered on bringing a sense of clarity and insight to complex topics ranging from the regulation of emerging technologies to the esoteric ideologies of Silicon Valley executives, while striving not to lose the poetic sense of awe inspired by often-obscure fields like astrophysics and quantum computing. I broke the story of CNET using AI to produce articles that turned out to be riddled with factual errors and plagiarism — a dam-breaking inflection point, as I've reported, that's inspired copycats and endless discourse while beguiling stakeholders ranging from tech giants to purveyors of spam around the web. My work at Futurism has been cited by publications including CBS News, the Los Angeles Times, Vice, Gizmodo, Engadget, the Verge, and Vanity Fair. I grew up in locales ranging from India to China, and now live in the exotic suburbs of Virginia. In my free time, I'm an avid reader of weird sci-fi literature, an aficionado of East Asian cinema, and, regrettably, a relapsed gamer. Allegedly, I’m working on a debut novel, currently untitled.

### Author social links  
[Bluesky](<https://bsky.app/profile/f-w-l.bsky.social>)