---
title: "OpenAI Research Finds That Even Its Best Models Give Wrong Answers a Wild Proportion of the Time"
description: "OpenAI has released a new benchmark dubbed \"SimpleQA\" to measure the accuracy of its AI models. The results are damning."
date: "2024-11-02"
modified: "2024-11-02"
authors:
  - name: "Victor Tangermann"
    job_title: "Senior Editor"
    link: "https://futurism.com/authors/victor"
url: "https://futurism.com/the-byte/openai-research-best-models-wrong-answers"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai chatbots"
  - "chatgpt"
  - "large language models"
  - "o1-preview"
  - "OpenAI"
  - "the digest"
---

# OpenAI Research Finds That Even Its Best Models Give Wrong Answers a Wild Proportion of the Time

![OpenAI has released a new benchmark dubbed "SimpleQA" to measure the accuracy of its AI models. The results are damning.](<https://futurism.com/wp-content/uploads/2024/10/openai-research-best-models-wrong-answers.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## BS Generator

OpenAI has [released a new benchmark](<https://openai.com/index/introducing-simpleqa/>), dubbed "SimpleQA," that's designed to measure the accuracy of the output of its own and competing artificial intelligence models.

In doing so, the AI company has revealed just how bad its latest models are at providing correct answers. In its own tests, its cutting edge o1-preview model, which was [released last month](<https://futurism.com/openai-released-strawberry-o1-preview-model>), scored an abysmal 42.7 percent success rate on the new benchmark.

In other words, even the cream of the crop of recently announced large language models (LLMs) is far more likely to provide an outright incorrect answer than a right one — a concerning indictment, especially as the tech is starting to pervade many aspects of our everyday lives.

## Wrong Again

Competing models, like Anthropic's, scored even lower on OpenAI's SimpleQA benchmark, with its recently released Claude-3.5-sonnet model getting only 28.9 percent of questions right. However, the model was far more inclined to reveal its own uncertainty and decline to answer — which, given the damning results, is probably for the best.

Worse yet, OpenAI found that its own AI models tend to vastly overestimate their own abilities, a characteristic that can lead to them being highly confident in the falsehoods they concoct.

LLMs have long suffered from "hallucinations," an elegant term AI companies have come up with to denote their models' [well-documented tendency](<https://futurism.com/perplexity-webpage-mushrooms>) to produce answers that are complete BS.

Despite the very high chance of ending up with complete fabrications, the world has embraced the tech with open arms, from students [generating homework assignments](<https://futurism.com/the-byte/students-admit-chatgpt-homework>) to developers employed by tech giants generating [huge swathes of code](<https://futurism.com/the-byte/google-ceo-code-ai>).

And the cracks are starting the show. Case in point, an AI model used by hospitals and built on OpenAI tech was [caught this week](<https://futurism.com/the-byte/whisper-nabla-hospital-ai-details-patients>) introducing frequent hallucinations and inaccuracies while transcribing patient interactions.

Cops across the United States are also [starting to embrace AI](<https://futurism.com/hallucinating-ai-police-reports>), a terrifying development that could lead to law enforcement falsely accusing the innocent or furthering troubling biases.

OpenAI's latest findings are yet another worrying sign that current LLMs are woefully unable to reliably tell the truth.

It's a development that should serve as a reminder to treat any output of any LLM out there with plenty of skepticism and a willingness to go over the generated text with a fine-toothed comb.

Whether it's a problem that can be solved with even bigger training sets — something AI leaders are [rushing to assure investors](<https://fortune.com/2024/06/06/ai-training-bottleneck-google-meta-openai/>) of — remains an [open question](<https://futurism.com/the-byte/ceo-google-ai-hallucinations>).

**More on OpenAI:** *[AI Model Used By Hospitals Caught Making Up Details About Patients, Inventing Nonexistent Medications and Sexual Acts](<https://futurism.com/the-byte/whisper-nabla-hospital-ai-details-patients>)*

## Author
I've been at Futurism since 2017, where my role has evolved to encompass design, writing, and increasingly editing. I've always been fascinated by space exploration and advanced transportation, which I've leaned into by interviewing luminaries in those fields while closely following the dimensions of policy and regulation that allow next-generation projects to succeed -- or, sometimes, to fail. I'm also keenly interested in the effects of generative AI on society, policies, and democratic institutions, as well as clean energy, physics and biology, and the vagaries of tech leadership. My work for Futurism has been cited by publications including Ars Technica, Gizmodo, PC Magazine, Jalopnik, Fox News, and the New York Post. I spent my childhood living in locations including Manila, the Philippines, and Geneva, Switzerland, attended McGill University, and now live in Toronto, Canada. Before Futurism I worked at AskMen and a small photography studio. In my free time, I'm an avid gardener, foodie, and craft beer lover, as well as a maker of artisanal hot pepper sauces. I have a magnificent dog named Freida.

### Author social links  
[Bluesky](<https://bsky.app/profile/vtanger.bsky.social>)