---
title: "AI Systems Are Learning to Lie and Deceive, Scientists Find"
description: "AI models are, apparently, getting better at lying on purpose -- and that might not be good news for us humans."
date: "2024-06-07"
modified: "2024-06-07"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/ai-systems-lie-deceive"
categories:
  - "Artificial Intelligence"
tags:
  - "diplomacy"
  - "gpt-4"
  - "llms"
  - "Meta"
---

# AI Systems Are Learning to Lie and Deceive, Scientists Find

![AI models are, apparently, getting better at lying on purpose -- and that might not be good news for us humans.](<https://futurism.com/wp-content/uploads/2024/06/ai-systems-lie-deceive.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

AI models are, apparently, getting better at lying on purpose.

Two recent studies — one [published this week in the journal *PNAS*](<https://www.pnas.org/doi/full/10.1073/pnas.2317967121>) and the other [last month in the journal *Patterns*](<https://www.cell.com/action/showPdf?pii=S2666-3899%2824%2900103-X>) — reveal some jarring findings about large language models (LLMs) and their ability to lie to or deceive human observers on purpose.

In the *PNAS* paper, German AI ethicist Thilo Hagendorff goes so far as to say that sophisticated LLMs can be encouraged to elicit "Machiavellianism," or intentional and amoral manipulativeness, which "can trigger misaligned deceptive behavior."

"GPT- 4, for instance, exhibits deceptive behavior in simple test scenarios 99.16% of the time," the University of Stuttgart researcher writes, citing his own experiments in quantifying various "maladaptive" traits in 10 different LLMs, most of which are different versions within OpenAI's GPT family.

Billed as a [human-level champion](<https://ai.meta.com/research/cicero/diplomacy/>) in the political strategy board game "Diplomacy," Meta's Cicero model was the subject of the *Patterns* study. As the disparate research group — comprised of a physicist, a philosopher, and two AI safety experts — found, the LLM got ahead of its human competitors by, in a word, fibbing.

Led by Massachusetts Institute of Technology postdoctoral researcher Peter Park, that paper found that Cicero not only excels at deception, but seems to have learned how to lie the more it gets used — a state of affairs "much closer to explicit manipulation" than, say, [AI's propensity for hallucination](<https://futurism.com/the-byte/amazon-ai-severe-hallucinations>), in which models confidently assert the wrong answers accidentally.

While Hagendorff notes in his more recent paper that the issue of LLM deception and lying is confounded by AI's inability to have any sort of human-like "intention" in the human sense, the *Patterns* study argues that within the confines of Diplomacy, at least, Cicero seems to break its programmers' promise that the model will "never intentionally backstab" its game allies.

The model, as the older paper's authors observed, "engages in premeditated deception, breaks the deals to which it had agreed, and tells outright falsehoods."

Put another way, as Park explained in a press release: "We found that Meta’s AI had learned to be a master of deception."

"While Meta succeeded in training its AI to win in the game of Diplomacy," the MIT physicist said in the school's statement, "Meta failed to train its AI to win honestly."

In a [statement to the *New York Post*](<https://nypost.com/2024/05/14/business/metas-ai-system-cicero-beats-humans-in-game-of-diplomacy-by-lying-study/>) after the research was first published, Meta made a salient point when echoing Park's assertion about Cicero's manipulative prowess: that "the models our researchers built are trained solely to play the game Diplomacy."

Well-known for expressly allowing lying, Diplomacy has jokingly been referred to as a [friendship-ending game](<https://foreignpolicy.com/2020/10/23/the-game-that-ruins-friendships-and-shapes-careers/>) because it encourages pulling one over on opponents, and if Cicero was trained exclusively on its rulebook, then it was essentially trained to lie.

Reading between the lines, neither study has demonstrated that AI models are lying over their own volition, but instead doing so because they've either been trained or jailbroken to do so.

That's good news for those concerned about AI developing sentience — but very bad news if you're worried about someone building an LLM with mass manipulation as a goal.

**More on bad AI:** [*News Site Says It’s Using to AI to Crank Out Articles Bylined by Fake Racially Diverse Writers in a Very Responsible Way*](<https://futurism.com/ai-fake-racially-diverse-writers>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)