---
title: "ChatGPT Actually Gets Half of Programming Questions Wrong"
description: "Researchers found that ChatGPT got just half of 517 Stack Overflow prompts wrong, indicating it's not great at spitting out perfect code."
date: "2023-08-10"
modified: "2023-08-10"
authors:
  - name: "Victor Tangermann"
    job_title: "Senior Editor"
    link: "https://futurism.com/authors/victor"
url: "https://futurism.com/the-byte/chatgpt-half-programming-questions-wrong"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai chatbots"
  - "chatgpt"
  - "OpenAI"
  - "Programming"
  - "the digest"
---

# ChatGPT Actually Gets Half of Programming Questions Wrong

![Researchers found that ChatGPT got just half of 517 Stack Overflow prompts wrong, indicating it's not great at spitting out perfect code.](<https://futurism.com/wp-content/uploads/2023/08/chatgpt-half-programming-questions-wrong.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Failing Grade

Not long after it was released to the public, programmers started to take note of a notable feature of OpenAI's ChatGPT: that it could [quickly spit out code](<https://www.theguardian.com/commentisfree/2023/apr/01/chatgpt-write-code-computer-programmer-software>), in response to easy prompts.

But should software engineers really trust its output?

In a [yet-to-be-peer-reviewed study](<https://arxiv.org/pdf/2308.02312.pdf>), researchers at Purdue University found that the uber-popular AI tool got just over half of 517 software engineering prompts from the popular question-and-answer platform Stack Overflow wrong — a sobering reality check that should have programmers think twice before deploying ChatGPT's answers in anything important.

## Pathological Liar

The research goes further, though, finding intriguing nuance in the ability of humans as well. The researchers asked a group of 12 participants with varying levels of programming expertise to analyze ChatGPT's answers. While they tended to rate Stack Overflow's answers higher across categories including correctness, comprehensiveness, conciseness, and usefulness, they weren't great at identifying the answers ChatGPT got wrong, failing to identify incorrect answers 39.34 percent of the time.

In other words, ChatGPT is a very convincing liar — a reality we've become [all too familiar with](<https://futurism.com/the-byte/google-staff-warned-ai-pathological-liar>).

"Users overlook incorrect information in ChatGPT answers (39.34 percent of the time) due to the comprehensive, well-articulated, and humanoid insights in ChatGPT answers," the paper reads.

So how worried should we really be? For one, there are many ways to arrive at the same "correct" answer in software. A lot of human programmers also say they verify ChatGPT's output, suggesting they understand the tool's limitations. But whether that'll continue to be the case remains to be seen.

## Lack of Reason

The researchers argue that a lot of work still needs to be done to address these shortcomings.

"Although existing work focus on removing hallucinations from \[large language models\], those are only applicable to fixing factual errors," they write. "Since the root of conceptual error is not hallucinations, but rather a lack of understanding and reasoning, the existing fixes for hallucination are not applicable to reduce conceptual errors."

In response, we need to focus on "teaching ChatGPT to reason," the researchers conclude — a tall order for this current generation of AI.

**More on ChatGPT:** *[AI Expert Says ChatGPT Is Way Stupider Than People Realize](<https://futurism.com/the-byte/ai-expert-chatgpt-way-stupider>)*

## Author
I've been at Futurism since 2017, where my role has evolved to encompass design, writing, and increasingly editing. I've always been fascinated by space exploration and advanced transportation, which I've leaned into by interviewing luminaries in those fields while closely following the dimensions of policy and regulation that allow next-generation projects to succeed -- or, sometimes, to fail. I'm also keenly interested in the effects of generative AI on society, policies, and democratic institutions, as well as clean energy, physics and biology, and the vagaries of tech leadership. My work for Futurism has been cited by publications including Ars Technica, Gizmodo, PC Magazine, Jalopnik, Fox News, and the New York Post. I spent my childhood living in locations including Manila, the Philippines, and Geneva, Switzerland, attended McGill University, and now live in Toronto, Canada. Before Futurism I worked at AskMen and a small photography studio. In my free time, I'm an avid gardener, foodie, and craft beer lover, as well as a maker of artisanal hot pepper sauces. I have a magnificent dog named Freida.

### Author social links  
[Bluesky](<https://bsky.app/profile/vtanger.bsky.social>)