---
title: "Research Shows That AI-Generated Slop Overuses Specific Words"
description: "By analyzing a decade of scientific papers, researchers found AI models are overusing \"style\" words that were uncommon just a few years ago."
date: "2024-07-02"
modified: "2024-07-02"
authors:
  - name: "Frank Landymore"
    job_title: "Contributing Writer"
    link: "https://futurism.com/authors/flandymore"
url: "https://futurism.com/the-byte/ai-overuses-specific-words"
categories:
  - "Artificial Intelligence"
tags:
  - "chatgpt"
  - "generative ai"
  - "large language models"
  - "the digest"
---

# Research Shows That AI-Generated Slop Overuses Specific Words

![By analyzing a decade of scientific papers, researchers found AI models are overusing "style" words that were uncommon just a few years ago.](<https://futurism.com/wp-content/uploads/2024/07/ai-overuses-specific-words.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Disease Control

AI models may be trained on the [entire corpus of humanity's writing](<https://futurism.com/the-byte/microsoft-ceo-ai-open-web>), but it turns out their vocabulary can be strikingly limited. A new [yet-to-be-peer-reviewed study](<https://arxiv.org/pdf/2406.07016>), spotted [by *Ars Technica*](<https://arstechnica.com/ai/2024/07/the-telltale-words-that-could-identify-generative-ai-text>), adds to the general understanding that large language models tend to overuse certain words [that can give their origins away](<https://www.washingtonpost.com/technology/2024/01/20/openai-use-policy-ai-writing-amazon-x/>).

In a novel approach, these researchers took a cue from epidemiology by measuring "excess word usage" in biomedical papers in the same way doctors gauged COVID-19's impact through "excess deaths." The results are a fascinating insight into AI's impact in the world of academia, suggesting that at least 10 percent of abstracts in 2024 were "processed with LLMs."

"The effect of LLM usage on scientific writing is truly unprecedented and outshines even the drastic changes in vocabulary induced by the COVID-19 pandemic," the researchers wrote in the study.

The work may even provide a boost for methods of detecting AI writing, which have so far proved [notoriously unreliable](<https://futurism.com/ai-plagiarism-software-false-accusing-students>).

## Style Over Substance

These findings come from a broad analysis of 14 million biomedical abstracts published between 2010 and 2024 that are available on PubMed. The researchers used papers published before 2023 as a baseline to compare papers that came out during the widespread commercialization of LLMs like ChatGPT.

They found that words that were once considered "less common," like "[delves](<https://www.businessinsider.com/y-combinator-paul-graham-delve-ai-chatgpt-giveaway-email-pitch-2024-4>)," are now used 25 more times than they used to, and others, like "showcasing" and "underscores," saw a similarly baffling nine times increase. But some "common" words also saw a boost: "potential," "findings," and "crucial" went up in frequency by up to 4 percent.

Such a marked increase is basically unprecedented without the explanation of some pressing global circumstance. When the researchers looked for excess words between 2013 and 2023, the ones that came up were terms like "ebola," "coronavirus," and "lockdown."

Beyond their obvious ties to real-world events, these are all nouns, or as the researchers put it, "content" words. By contrast, what we see with the excess usage in 2024 is that they're almost entirely "style" words. And in numbers, of the 280 excess "style" words that year, two-thirds of them were verbs, and about a fifth were adjectives.

To see just how saturated AI language is with these tell-tales, have a look at this example from a real 2023 paper (emphasis the researchers'): "By **meticulously delving** into the **intricate** web connecting \[...\] and \[...\], this **comprehensive** chapter takes a deep dive into their involvement as **significant** risk factors for \[...\].

## Language Barriers

Using these excess style words as "markers" of ChatGPT usage, the researchers estimated that around 15 percent of papers published in non-English speaking countries like China, South Korea, and Taiwan are now AI-processed — which is higher than in countries where English is the native tongue, like the United Kingdom, at 3 percent. LLMs, then, may be a genuinely helpful tool for non-native speakers to make it in a field [dominated by English](<https://theconversation.com/prestigious-journals-make-it-hard-for-scientists-who-dont-speak-english-to-get-published-and-we-all-lose-out-226225>).

Still, the researchers admit that native speakers may simply be better at hiding their LLM usage. And of course, the appearance of these words is not a guarantee that the text was AI-generated.

Whether this will serve as a reliable detection method is up in the air — but what is certainly evidence here is just how quickly AI can catalyze changes in written language.

**More on AI:** *[AI Researcher Elon Musk Poached From OpenAI Returns to OpenAI](<https://futurism.com/the-byte/xai-researcher-returns-openai>)*

## Author
At Futurism, my work has often centered on bringing a sense of clarity and insight to complex topics ranging from the regulation of emerging technologies to the esoteric ideologies of Silicon Valley executives, while striving not to lose the poetic sense of awe inspired by often-obscure fields like astrophysics and quantum computing. I broke the story of CNET using AI to produce articles that turned out to be riddled with factual errors and plagiarism — a dam-breaking inflection point, as I've reported, that's inspired copycats and endless discourse while beguiling stakeholders ranging from tech giants to purveyors of spam around the web. My work at Futurism has been cited by publications including CBS News, the Los Angeles Times, Vice, Gizmodo, Engadget, the Verge, and Vanity Fair. I grew up in locales ranging from India to China, and now live in the exotic suburbs of Virginia. In my free time, I'm an avid reader of weird sci-fi literature, an aficionado of East Asian cinema, and, regrettably, a relapsed gamer. Allegedly, I’m working on a debut novel, currently untitled.

### Author social links  
[Bluesky](<https://bsky.app/profile/f-w-l.bsky.social>)