---
title: "Stanford Scientists Find That Yes, ChatGPT Is Getting Stupider"
description: "AI Inaccuracy While looking into artificial intelligence \"behavior,\" researchers affirmed that yes, OpenAI's GPT large language model (LLM) appeared to be getting dumber. In a pre-print study, researchers out of Stanford and Berkeley found that over a period of a few months, both GPT-4.5 and GPT-4 significantly changed their \"behavior\" and that the accuracy of […]"
date: "2023-07-20"
modified: "2023-07-20"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/stanford-chatgpt-getting-dumber"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "chatgpt"
  - "gpt-4"
  - "OpenAI"
  - "the digest"
---

# Stanford Scientists Find That Yes, ChatGPT Is Getting Stupider

![While looking into artificial intelligence "behavior," researchers affirmed that yes, OpenAI's GPT-4 appeared to be getting dumber. ](<https://futurism.com/wp-content/uploads/2023/07/stanford-chatgpt-getting-dumber.jpg>)
*Head in the sand: white robots kneeling with head buried in the ground \<em\>Image: Getty Images\</em\>*

## Dumb and Dumber

Regardless of [what its execs claim](<https://futurism.com/the-byte/openai-denies-chatgpt-stupider>), researchers are now saying that yes, OpenAI's GPT large language model (LLM) appeared to be getting dumber.

In a [new yet-to-be-peer-reviewed study](<https://arxiv.org/pdf/2307.09009.pdf>), researchers out of Stanford and Berkeley found that over a period of a few months, both GPT-3.5 and GPT-4 significantly changed their "behavior," with the accuracy of their responses appearing to go down, validating user anecdotes about the apparent degradation of the latest versions of the software in the months since their releases.

"GPT-4 (March 2023) was very good at identifying prime numbers (accuracy 97.6 percent)," the researchers wrote in their paper's abstract, "but GPT-4 (June 2023) was very poor on these same questions (accuracy 2.4 percent)."

"Both GPT-4 and GPT-3.5," the abstract continued, "had more formatting mistakes in code generation in June than in March."

## Brain Drain

This study affirms what users have been saying for more than a month now: that as they've used the GPT-3 and GPT-4-powered ChatGPTover time, they've [noticed it becoming, well, stupider](<https://www.businessinsider.com/openai-gpt4-ai-model-got-lazier-dumber-chatgpt-2023-7>).

The seeming degradation of its accuracy has become so troublesome that OpenAI vice president of product Peter Welinder attempted to dispel rumors that the change was intentional.

"No, we haven't made GPT-4 dumber," [Welinder tweeted last week](<https://twitter.com/npew/status/1679538687854661637>). "Quite the opposite: we make each new version smarter than the previous one."

He added that changes in user experience could be due to continuous use, saying that it could be that "when you use \[ChatGPT\] more heavily, you start noticing issues you didn't see before."

## Class Clown

The Stanford and Berkeley research is a compelling datapoint against that hypothesis, though. While the researchers don't posit reasons as to why these downward "drifts" in accuracy and ability are occurring, they do note that this demonstrable worsening over time challenges OpenAI's insistence that its models are instead improving.

"We find that the performance and behavior of both GPT-3.5 and GPT-4 vary significantly across these two releases and that their performance on some tasks have gotten substantially worse over time," the paper noted, adding that it's "interesting" to question whether GPT-4 is indeed getting stronger.

"It’s important to know whether updates to the model aimed at improving some aspects actually hurt its capability in other dimensions," the researchers wrote.

Translation: OpenAI's rapid updates may be doing more harm than good for ChatGPT, which has already become [known for its inaccuracies](<https://futurism.com/insider-chatgpt-plans>).

**More on OpenAI:** [*Theory: ChatGPT Use Is Falling Because Kids Don't Need to Cheat on Papers During Summer Vacation*](<https://futurism.com/the-byte/theory-chatgpt-students-cheating-summer>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)