---
title: "OpenAI Admits That Its New Model Still Hallucinates More Than a Third of the Time"
description: "If a partner made stuff up a third of the time, it would be a problem — but apparently, it's perfectly fine for OpenAI's new model."
date: "2025-03-01"
modified: "2025-03-01"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/openai-admits-gpt45-hallucinates"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai hallucinations"
  - "gpt-4.5"
  - "llms"
  - "OpenAI"
---

# OpenAI Admits That Its New Model Still Hallucinates More Than a Third of the Time

![If a partner made stuff up a third of the time, it would be a problem — but apparently, it's perfectly fine for OpenAI's new model.](<https://futurism.com/wp-content/uploads/2025/02/openai-admits-gpt45-hallucinates.jpg>)
*\<em\>Image: Joel Saget / AFP via Getty / Futurism\</em\>*

If a partner or friend made stuff up a significant percentage of the time that you asked a question, it would be a huge problem for the relationship.

But apparently it's different for OpenAI's hot new model. Using SimpleQA, the company's in-house [factuality benchmarking tool](<https://openai.com/index/introducing-simpleqa/>), OpenAI admitted in its release announcement that its new large language model (LLM) GPT-4.5 hallucinates — which is AI parlance for confidently spewing fabrications and presenting them as fact — 37 percent of the time.

Yes, you read that right: in tests, the latest AI model from a company that's [worth hundreds of billions of dollars](<https://www.nytimes.com/2025/02/07/technology/openai-softbank-investment.html>) is telling lies for more than one out of every three answers it gives.

As if that wasn't bad enough, OpenAI is actually trying to spin GPT-4.5's bullshitting problem as a *good* thing because — get this — it doesn't hallucinate as much as the company's other LLMs.

The same graph \[can we embed a screenshot below?\] that showed how often the new model spews nonsense also reports that GPT-4o, a purportedly advanced "reasoning" model, hallucinates 61.8 percent of the time on the SimpleQA benchmark. OpenAI's o3-mini, a cheaper and smaller version of its reasoning model, was found to hallucinate a whopping 80.3 percent of the time.

Of course, the problem isn't unique to OpenAI.

"At present, even the best models can generate hallucination-free text only about 35 percent of the time," explained Wenting Zhao, a Cornell doctoral student who co-wrote a [paper last year](<https://www.arxiv.org/pdf/2407.17468>) about AI hallucination rates, in an [interview about the research with *TechCrunch*](<https://techcrunch.com/2024/08/14/study-suggests-that-even-the-best-ai-models-hallucinate-a-bunch/>). "The most important takeaway from our work is that we cannot yet fully trust the outputs of model generations."

Beyond the incredulity of a company getting hundreds of billions of dollars in investments for products that have such issues telling the truth, it says a lot about the AI industry at large that these are the things they're selling us: expensive, resource-consuming systems that are supposed to be approaching human-level intelligence but still can't get basic facts right.

As [OpenAI's LLMs plateau](<https://www.axios.com/2024/11/13/ai-scaling-chatgpt-openai-plateau>) in performance, the company is clearly grasping at straws to re-steer the hype ship back on the [course it seemed to chart](<https://fortune.com/2025/02/25/what-happened-gpt-5-openai-orion-pivot-scaling-pre-training-llm-agi-reasoning/>) when ChatGPT first dropped.

But to do that, we're probably going to need to see a real breakthrough, not more of the same.

**More on AI hallucinations:** [*Even the Most Advanced AI Has a Problem: If It Doesn’t Know the Answer, It Makes One Up*](<https://futurism.com/ai-makes-up-answers>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)