---
title: "GPT-5 Is Making Huge Factual Errors, Users Say"
description: "OpenAI doesn't seem to have fixed GPT-5's problems with truth and accuracy in the month since the LLM was released."
date: "2025-09-09"
modified: "2025-09-09"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/gpt-5-huge-factual-errors"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai hallucinations"
  - "chatgpt"
  - "gpt-5"
  - "OpenAI"
---

# GPT-5 Is Making Huge Factual Errors, Users Say

![OpenAI doesn't seem to have fixed GPT-5's problems with truth and accuracy in the month since the LLM was released.](<https://futurism.com/wp-content/uploads/2025/09/gpt-5-huge-factual-errors.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

It's been just over a month since OpenAI dropped its [long-awaited GPT-5](<https://futurism.com/openai-releases-gpt-5>) large language model (LLM) — and it hasn't stopped spewing an astonishing amount of strange falsehoods since then.

From the [AI experts](<https://mindmatters.ai/2025/09/gpt-5-0-doesnt-understand-but-is-eager-to-please/>) at the Discovery Institute's Walter Bradley Center for Artificial Intelligence and irked Redditors on r/ChatGPTPro, to even OpenAI CEO [Sam Altman himself](<https://futurism.com/sam-altman-admits-openai-screwed-up>), there's plenty of evidence to suggest that OpenAI's claim that GPT-5 boasts "[PhD-level intelligence](<https://openai.com/index/introducing-gpt-5/>)" comes with some serious asterisks.

In a Reddit post, a [user realized](<https://www.reddit.com/r/ChatGPTPro/comments/1n890r6/chatgpt_5_has_become_unreliable_getting_basic/>) not only that GPT-5 had been generating "wrong information on basic facts over half the time," but that without fact-checking, they may have missed other hallucinations.

The Reddit user's experience highlights just how common it is for chatbots to hallucinate, which is AI-speak for confidently making stuff up. While the issue is [far from exclusive to ChatGPT](<https://futurism.com/ai-industry-problem-smarter-hallucinating>), OpenAI's latest LLM seems to have a [particular penchant for BS](<https://futurism.com/the-byte/researchers-ai-chatgpt-hallucinations-terminology>) — a reality that challenges the company's claim that [GPT-5 hallucinates less](<https://mashable.com/article/openai-gpt-5-hallucinates-less-system-card-data>) than its predecessors.

In a [recent blog post about hallucinations](<https://openai.com/index/why-language-models-hallucinate/>), in which OpenAI once again claimed that GPT-5 produces "significantly fewer" of them — the firm attempted to explain how and why these falsehoods occur.

"Hallucinations persist partly because current evaluation methods set the wrong incentives," the September 5 post reads. "While evaluations themselves do not directly cause hallucinations, most evaluations measure model performance in a way that encourages guessing rather than honesty about uncertainty."

Translation: LLMs hallucinate because they are trained to get things right, even if it means guessing. Though some models, [like Anthropic's Claude](<https://futurism.com/ai-makes-up-answers>), have been trained to admit when they don't know an answer, OpenAI's have not — thus, they wager incorrect guesses.

As the Reddit user indicated (backed up with a [link to their conversation log](<https://chatgpt.com/share/68b99a61-5d14-800f-b2e0-7cfd3e684f15>)), they got some massive factual errors when asking about the gross domestic product (GDP) of various countries and were presented by the chatbot with "figures that were literally double the actual values."

Poland, for instance, was listed as having a GDP of more than two *trillion* dollars, when in reality its GDP, [per the International Monetary Fund](<https://www.imf.org/external/datamapper/profile/POL>), is currently hovering around $979 billion. Were we to wager a guess, we'd say that that hallucination may be attributed to [recent boasts from the country's president](<https://www.gov.pl/web/primeminister/poland-joins-the-trillionaires-club-a-historic-entry-into-the-worlds-top-20-economies#:~:text=Poland%20Among%20Economic%20Leaders,exclusive%20club%20of%20trillionaire%20countries.>) saying its economy (and not its GDP) has exceeded $1 trillion.

"The scary part? I only noticed these errors because some answers seemed so off that they made me suspicious," the user continued. "For instance, when I saw GDP numbers that seemed way too high, I double-checked and found they were completely wrong."

"This makes me wonder: How many times do I NOT fact-check and just accept the wrong information as truth?" they mused.

Meanwhile, AI skeptic Gary Smith of the Walter Bradley Center noted that he's done three simple experiments with GPT-5 since its release — a [modified game of tic-tac-toe](<https://futurism.com/gpt-5-simple-question-confusion>), [questioning about financial advice](<https://mindmatters.ai/2025/08/what-kind-of-a-phd-level-expert-is-chatgpt-5-0-i-tested-it/>), and a [request to draw a possum](<https://mindmatters.ai/2025/08/what-kind-of-a-phd-level-expert-is-chatgpt-5-0-i-tested-it/>) with five of its body parts labeled — to "demonstrate that GPT 5.0 was far from PhD-level expertise."

The [possum example](<https://mindmatters.ai/2025/08/what-kind-of-a-phd-level-expert-is-chatgpt-5-0-i-tested-it/>) was particularly egregious, technically coming up with the right names for the animal's parts but pinning them in strange places, such as marking its leg as its nose and its tail as its back left foot. When attempting to replicate the experiment for a [more recent post](<https://mindmatters.ai/2025/09/gpt-5-0-doesnt-understand-but-is-eager-to-please/>), Smith discovered that even when he made a typo — "posse" instead of "possum" — GPT-5 mislabeled the parts in a similarly bizarre fashion.

Instead of the intended possum, the LLM generated an image of its apparent idea of a posse: five cowboys, some toting guns, with lines indicating various parts. Some of those parts — the head, foot, and possibly the ear — were accurate, while the shoulder pointed to one of the cowboys' ten-gallon hats and the "fand," which may be a mix-up of foot and hand, pointed at one of their shins.

We decided to do a similar test, asking GPT-5 to provide an image of "[a posse with six body parts labeled](<https://chatgpt.com/share/68c04916-1618-800f-b396-cb0cde072a02>)." After clarifying that *Futurism* wanted a labeled image and not a text description, ChatGPT went off to work — and what it spat out was, as you can see below, even more hilariously wrong than what Smith got.

![A bizarre GPT-5-generated image of a six-person cowboy posse with strangely-named body parts labeled. Up top, the second on the left cowboy's hat is labeled "head," while the second-to-last's hat is labeled "mouth." On bottom, each of the six gunslingers — two men, a woman, and three other men from left to right — have bizarre labels beneath them that read "harn, old, hand, hag, leg -\> spur." Image via ChatGPT/Futurism.](<https://futurism.com/wp-content/uploads/2025/09/Old-West-Posse-with-Labeled-Details-2.png>)

It seems pretty clear from this side of the GPT-5 release that it's nowhere near as smart as a doctoral candidate — or, at very least, one that has any chance of actually attaining their PhD.

The moral of this story, it seems, is to fact-check anything a chatbot spits out — or forgo using AI and do the research for yourself.

**More on GPT-5:** [*After Disastrous GPT-5, Sam Altman Pivots to Hyping Up GPT-6*](<https://futurism.com/disastrous-gpt-5-sam-altman-hyping-up-gpt-6>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)