---
title: "OpenAI’s Hot New AI Has an Embarrassing Problem"
description: "OpenAI's latest AI models tend to make things up — or \"hallucinate\" — substantially more than earlier versions."
date: "2025-04-21"
modified: "2025-04-21"
authors:
  - name: "Victor Tangermann"
    job_title: "Senior Editor"
    link: "https://futurism.com/authors/victor"
url: "https://futurism.com/the-byte/openai-new-ai-problem-hallucinate-more"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "chatgpt"
  - "generative ai"
  - "OpenAI"
  - "the digest"
---

# OpenAI’s Hot New AI Has an Embarrassing Problem

![OpenAI's latest AI models tend to make things up — or "hallucinate" — substantially more than earlier versions.](<https://futurism.com/wp-content/uploads/2025/04/openai-new-ai-problem-hallucinate-more.jpg>)
*\<em\>Image: Sean Gallup/Getty Images\</em\>*

## Bucking the Trend

OpenAI [launched](<https://openai.com/index/o3-o4-mini-system-card/>) its latest AI reasoning models, dubbed o3 and o4-mini, last week.

According to the Sam Altman-led company, the new models outperform their predecessors and "excel at solving complex math, coding, and scientific challenges while demonstrating strong visual perception and analysis."

But there's one extremely important area where o3 and o4-mini appear to instead be taking a major step back: they tend to make things up — or "hallucinate" — substantially *more* than those earlier versions, as [*TechCrunch* reports](<https://techcrunch.com/2025/04/18/openais-new-reasoning-ai-models-hallucinate-more/>).

The news once again highlights a nagging technical issue that has plagued the industry [for years now](<https://futurism.com/the-byte/google-gmail-tool-hallucinating-emails>). Tech companies have struggled to rein in rampant hallucinations, which have greatly [undercut the usefulness](<https://futurism.com/amazon-ai-powered-alexa-hallucinations-problem>) of tools like ChatGPT.

Worryingly, OpenAI's two new models also buck a historical trend, which has seen each new model incrementally hallucinating less than the previous one, as *TechCrunch* points out, suggesting OpenAI is now headed in the wrong direction.

## Head in the Clouds

According to OpenAI's own internal testing, o3 and o4-mini tend to hallucinate more than older models, including o1, o1-mini, and even o3-mini, which was released in late January.

Worse yet, the firm doesn't appear to fully understand why. According to its [technical report](<https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf>), "more research is needed to understand the cause" of the rampant hallucinations.

Its o3 model scored a hallucination rate of 33 percent on the company's in-house accuracy benchmark, dubbed PersonQA. That's roughly double the rate compared to the company's preceding reasoning models.

Its o4-mini scored an abysmal hallucination rate of 48 percent, part of which could be due to it being a smaller model that has "less world knowledge" and therefore tends to "hallucinate more," according to OpenAI.

Nonprofit AI research company Transluce also [found in its own testing](<https://transluce.org/investigating-o3-truthfulness>) that o3 had a strong tendency to hallucinate, especially when generating computer code.

The extent to which it tried to cover for its own shortcomings is baffling.

"It further justifies hallucinated outputs when questioned by the user, even claiming that it uses an external MacBook Pro to perform computations and copies the outputs into ChatGPT," Transluce wrote in its blog post.

Experts even told *TechCrunch* OpenAI's o3 model hallucinates broken website links that simply don't work when the user tries to click them.

Unsurprisingly, OpenAI is well aware of these shortcomings.

"Addressing hallucinations across all our models is an ongoing area of research, and we’re continually working to improve their accuracy and reliability," OpenAI spokesperson Niko Felix told *TechCrunch*.

**More on OpenAI:** *[OpenAI Is Secretly Building a Social Network](<https://futurism.com/openai-building-social-network>)*

## Author
I've been at Futurism since 2017, where my role has evolved to encompass design, writing, and increasingly editing. I've always been fascinated by space exploration and advanced transportation, which I've leaned into by interviewing luminaries in those fields while closely following the dimensions of policy and regulation that allow next-generation projects to succeed -- or, sometimes, to fail. I'm also keenly interested in the effects of generative AI on society, policies, and democratic institutions, as well as clean energy, physics and biology, and the vagaries of tech leadership. My work for Futurism has been cited by publications including Ars Technica, Gizmodo, PC Magazine, Jalopnik, Fox News, and the New York Post. I spent my childhood living in locations including Manila, the Philippines, and Geneva, Switzerland, attended McGill University, and now live in Toronto, Canada. Before Futurism I worked at AskMen and a small photography studio. In my free time, I'm an avid gardener, foodie, and craft beer lover, as well as a maker of artisanal hot pepper sauces. I have a magnificent dog named Freida.

### Author social links  
[Bluesky](<https://bsky.app/profile/vtanger.bsky.social>)