---
title: "Majority of Humans Fooled by GPT-4 in Turing Test, Scientists Find"
description: "OpenAI's GPT-4 is so lifelike, it can trick more than 50 percent of human test subjects into thinking they are talking to a person."
date: "2024-05-17"
modified: "2024-05-17"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/gpt-4-passed-turing-test"
categories:
  - "Artificial Intelligence"
tags:
  - "gpt-4"
  - "the digest"
  - "turing test"
  - "vitalik buterin"
---

# Majority of Humans Fooled by GPT-4 in Turing Test, Scientists Find

![OpenAI's GPT-4 is so lifelike, it can trick more than 50 percent of human test subjects into thinking they are talking to a person.](<https://futurism.com/wp-content/uploads/2024/05/gpt-4-passed-turing-test.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Pass/Fail

OpenAI's GPT-4 is so lifelike, it can apparently trick more than 50 percent of human test subjects into thinking they're talking to a person.

In a [new paper](<https://arxiv.org/abs/2405.08007>), cognitive science researchers from the University of California San Diego found that more than half the time, people mistook writing from GPT-4 as having been written by a flesh-and-blood human. In other words, the large language model (LLM) passes the Turing test with flying colors.

The researchers performed a simple experiment: they asked roughly 500 people to have five-minute text-based conversations with either a human or a chatbot built on GPT-4. They then asked the subjects if they thought they'd been conversing with a person or an AI.

The results, as the San Diego scientists reported in their not-yet-peer-reviewed paper, were telling: 54 percent of the subjects believed they'd been speaking to humans when they'd actually been chatting with OpenAI's creation.

First theorized back in 1950 by computer science pioneer Alan Turing, the [Turing Test](<https://plato.stanford.edu/entries/turing-test/>) is more of a thought experiment than an actual battery of tests. In his original test, Turing had three "players" — a human interrogator, a witness of indeterminate humanity or machine-ness, and a human observer.

For their study, the UC San Diego researchers tweaked Turing's original three-player formula by eliminating the third human observer to simplify the setup. They then had the 500 participants communicate with one of four witness types: another human, GPT-3.5, GPT-4, or the [rudimentary ELIZA chatbot](<https://web.njit.edu/~ronkowit/eliza.html>) from the 1960s.

## Coin Toss

Jones and Bergen hypothesized that the study's subjects would generally be able to tell most of the time if they were communicating with either a human or ELIZA, but that when it came to the OpenAI LLMs, they would essentially have a 50/50 chance.

As it turns out, they were pretty much on the money. Beyond the 54 percent who mistook GPT-4 for a human, exactly 50 percent of the subjects confused GPT-3.5, the latest LLM's direct predecessor, for a person as well. Compared to the 22 percent who thought ELIZA was the real deal, that's pretty stunning.

https://twitter.com/emollick/status/1790877242525942156

Despite still being under review, the paper has already made waves in the tech world with a [shoutout from Ethereum cofounder Vitalik Buterin](<https://warpcast.com/vitalik.eth/0xb12ba0c1>), who declared on the Farcaster social network that to his mind, the San Diego research "counts as \[GPT-4\] passing the Turing test."

While others have claimed to observe OpenAI's [GPT models passing the Turing test](<https://humsci.stanford.edu/feature/study-finds-chatgpts-latest-bot-behaves-humans-only-better>), the Buterin endorsement makes this study stand apart — though we'll probably have to wait for the paper to be peer-reviewed until any grander declarations can be made.

**More on GPT-4:** [*OpenAI Secretly Trained GPT-4 With More Than a Million Hours of Transcribed YouTube Videos*](<https://futurism.com/openai-gpt4-youtube>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)