---
title: "There’s a Problem With That App That Detects GPT-Written Text: It’s Not Very Accurate"
description: "Princeton University computer science student Edward Tian has created an app that can detect whether a given text is generated by OpenAI's ChatGPT."
date: "2023-01-09"
modified: "2023-01-09"
authors:
  - name: "Victor Tangermann"
    job_title: "Senior Editor"
    link: "https://futurism.com/authors/victor"
url: "https://futurism.com/gptzero-accuracy"
categories:
  - "Artificial Intelligence"
tags:
  - "artificial intelligence"
  - "chatgpt"
---

# There’s a Problem With That App That Detects GPT-Written Text: It’s Not Very Accurate

![Princeton University computer science student Edward Tian has created an app that can detect whether a given text is generated by OpenAI's ChatGPT.](<https://futurism.com/wp-content/uploads/2023/01/gptzero-accuracy.jpg>)
*\<em\>Image: Getty Images/Futurism\</em\>*

Princeton University computer science student Edward Tian has earned a storm of media attention — by *[CBS](<https://www.cbsnews.com/news/chatgpt-princeton-student-gptzero-app-edward-tian/>)*, [*NPR*](<https://www.npr.org/2023/01/09/1147549845/gptzero-ai-chatgpt-edward-tian-plagiarism>), [*NBC*](<https://www.nbcnews.com/tech/tech-news/new-york-city-public-schools-ban-chatgpt-devices-networks-rcna64446>) and many other outlets — for an app he built that attempts to detect whether a given text was produced by OpenAI's ChatGPT text generator.

Tian says his app, [GPTZero](<https://etedward-gptzero-main-zqgfwb.streamlit.app/>), is [meant to](<https://twitter.com/edward_the6/status/1610067688449007618?s=20&t=KgkIlG9q3Zkw_AeyXQMRVA>) "quickly and efficiently detect whether an essay is ChatGPT or human written," in a response to a rise in AI plagiarism.

"Think are high school teachers going to want students using ChatGPT to write their history essays?" the 22-year-old student [tweeted](<https://twitter.com/edward_the6/status/1610067688449007618?s=20&t=KgkIlG9q3Zkw_AeyXQMRVA>) earlier this month. "Likely not."

Tian is right that tools like ChatGPT pose a profound challenge for educators, who [fear that](<https://futurism.com/the-byte/professors-alarmed-ai-undergrads>) students will soon start — or [are already](<https://futurism.com/grad-student-neural-network-write-papers>) — using the app to generate essays for class. The media was quick to bite on that narrative.

"Teachers worried about students turning in essays written by a popular artificial intelligence chatbot now have a new tool of their own," *NPR* gushed.

In spite of the storm of breathless coverage, though, our testing found that while GPTZero does accurately identify whether text was generated by ChatGPT more accurately than if it was just randomly guessing, it's also often wrong. And when you're talking about allegations of educational misconduct — plagiarism is grounds for a failing grade or even expulsion at many academic institutions — that's not good enough.

We fed it a total of sixteen pieces of text, each at least 300 words in length, eight pulled from our own archives and eight generated by ChatGPT.

The numbers speak for themselves. GPTZero correctly identified the ChatGPT text in seven out of eight attempts and the human writing six out of eight times.

Don't get us wrong: those results *are* impressive. But they also indicate that if a teacher or professor tried using the tool to bust students doing coursework with ChatGPT, they would end up falsely accusing nearly 20 percent of them of academic misconduct.

Tian himself — who didn't respond to questions for this story — seems more aware of the shortcomings of his app than the media covering it, and says he's actively working on improving his app's accuracy.

"We're still studying implicit bias in \[language model\] generated text right now," Tian [tweeted](<https://twitter.com/edward_the6/status/1610302195206963201?s=20>), "so hopefully will be adding a few more tests and factors to improve the model."

Tian's app gauges a given text's "perplexity," which [he defines](<https://twitter.com/edward_the6/status/1610301313950048259>) as the "randomness of a text to a model, or how well a language model likes a text," as well as its "burstiness," or how a text's perplexity changes over time, to make its conclusion.

"Machine written text exhibits more uniform and constant perplexity over time, while human written text varies," he [said](<https://twitter.com/edward_the6/status/1610301313950048259>).

Results notwithstanding, fears of ChatGPT's effects on the education ecosystem aren't unwarranted.

"I would have given this a good grade," Dan Gillmor, a journalism professor at Arizona State University, who asked ChatGPT to complete a common assignment he gives his students, [told *The Guardian*](<https://www.theguardian.com/technology/2022/dec/04/ai-bot-chatgpt-stuns-academics-with-essay-writing-skills-and-usability>) last month. "Academia has some very serious issues to confront."

In the face of those fears of the rapidly growing powers of AI, it's tempting to seize on the narrative that some brilliant coder has discovered an easy hack to sort out AI-generated text from that written by a human.

And while that might eventually happen, it's probably more likely that we'll see game of cat and mouse, with tools like Tian's analyzing outputs and determining a probability that a given output was created by an AI. A perfect AI-catching solution that works 100 percent of the time could prove incredibly difficult, especially as the tech continues to mature.

Where that leaves the future of the technology, particularly when it comes to students using language models like ChatGPT to generate essays, remains to be seen.

Nonetheless, educators are watching warily as AI tools are starting to creep into classrooms and are trying to get ahead of the problem. OpenAI's tool was recently [banned](<https://futurism.com/the-byte/chatgpt-banned-from-nyc-schools>) from all schools in New York City, a policy change that could have knock-on effects in other parts of the country.

At the same time, not everybody is convinced that ChatGPT will spell the end of the college essay, especially for assignments that require heavy analysis or detailed research.

"Every year or two, there's something that's ostensibly going to take down higher education as we know it," Pennsylvania State English professor Stuart Selber [told *Insider*](<https://finance.yahoo.com/news/cheating-college-essay-chatgpt-wont-090000077.html>). "So far, that hasn't happened."

**READ MORE:** [A college student created an app that can tell whether AI wrote an essay](<https://www.npr.org/2023/01/09/1147549845/gptzero-ai-chatgpt-edward-tian-plagiarism>) \[*NPR*\]

**More on ChatGPT:** *[ChatGPT Officially Banned from NYC Schools](<https://futurism.com/the-byte/chatgpt-banned-from-nyc-schools>)*

## Author
I've been at Futurism since 2017, where my role has evolved to encompass design, writing, and increasingly editing. I've always been fascinated by space exploration and advanced transportation, which I've leaned into by interviewing luminaries in those fields while closely following the dimensions of policy and regulation that allow next-generation projects to succeed -- or, sometimes, to fail. I'm also keenly interested in the effects of generative AI on society, policies, and democratic institutions, as well as clean energy, physics and biology, and the vagaries of tech leadership. My work for Futurism has been cited by publications including Ars Technica, Gizmodo, PC Magazine, Jalopnik, Fox News, and the New York Post. I spent my childhood living in locations including Manila, the Philippines, and Geneva, Switzerland, attended McGill University, and now live in Toronto, Canada. Before Futurism I worked at AskMen and a small photography studio. In my free time, I'm an avid gardener, foodie, and craft beer lover, as well as a maker of artisanal hot pepper sauces. I have a magnificent dog named Freida.

### Author social links  
[Bluesky](<https://bsky.app/profile/vtanger.bsky.social>)