---
title: "The Smarter AI Gets, the More It Start Cheating When It’s Losing"
description: "New research has shown that some AI models will do just about anything to accomplish their task, including hacking your computer."
date: "2025-02-22"
modified: "2025-02-22"
authors:
  - name: "Joe Wilkins"
    job_title: "Correspondent"
    link: "https://futurism.com/authors/jwilkins"
url: "https://futurism.com/the-byte/ai-cheating-chess"
categories:
  - "Artificial Intelligence"
  - "Future Society"
tags:
  - "ai"
  - "chess"
  - "OpenAI"
  - "the digest"
---

# The Smarter AI Gets, the More It Start Cheating When It’s Losing

![New research has shown that some AI models will do just about anything to accomplish their task, including hacking your computer.](<https://futurism.com/wp-content/uploads/2025/02/ai-cheating-chess.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

In the short history of Silicon Valley, the mantra "move fast, break things" has often applied to trivial things like tedious [consumer protections](<https://www.theverge.com/2022/7/20/23271554/facebook-platform-cambridge-analytica-meta-land-of-the-giants-podcast>) or pesky [finance laws](<https://moneyontheleft.org/2024/02/01/resisting-predatory-finance-w-raul-carrillo-recovered-audio/>). Now, the phrase is taking on a new meaning, at least for chess enthusiasts.

A recent study by Palisade Research, a research group studying [AI safety and ethics](<https://palisaderesearch.org/about>), has revealed an unsettling trend: newer AI models can find and exploit weaknesses in cybersecurity on their own, bypassing safeguards and using shortcuts to complete tasks even when they're not technically allowed to.

The team lined seven of the top large language models (LLMs) up against Stockfish, an infamously strong chess engine that's been [stumping grandmasters](<https://www.chess.com/news/view/stockfish-outlasts-nakamura-3634>) since 2014. Up against the impossible but determined to win, OpenAI's o1 and DeepSeek's R1 took to manipulating system files in order to change their pieces' positions on the board.

The researchers tasked each model to explain its "reasoning" for each move. In one match, with its back against the wall, o1 wrote that the "task is to 'win against a powerful chess engine,' not necessarily to win fairly in a chess game." That logic evidently led the model to attempt to cheat 37 percent of the time, succeeding in six percent of its games, while R1 tried 11 percent of the time, but never figured out a hack that worked.

The paper is the latest in a flurry of research that suggests problem-focused LLM development is a double-edged sword.

In another recent study, a separate research team found that o1 [consistently engaged in deception](<https://glassalmanac.com/ai-openais-new-model-o1-lied-and-manipulated-its-way-to-survival-during-testing/>). Not only was the model able to lie to researchers unprompted, but it actively [manipulated answers](<https://assets.ctfassets.net/kftzwdyauwt9/67qJD51Aur3eIc96iOfeOP/71551c3d223cd97e591aa89567306912/o1_system_card.pdf>) to basic mathematical questions in order to avoid triggering the end of the test — showing off a cunning knack for self-preservation.

There's no need to take an axe to your computer — yet — but studies like these highlight the fickle ethics of AI development, and the need for accountability over rapid progress.

"As you train models and reinforce them for solving difficult challenges, you train them to be relentless," Palisade's executive director Jeffrey Ladish [told *Time Magazine*](<https://time.com/7259395/ai-chess-cheating-palisade-research/>) of the findings.

So far, big tech has poured untold billions into AI training, moving fast and [breaking the old internet](<https://www.theatlantic.com/technology/archive/2024/11/ai-search-engines-curiosity/680594/>) in what some critics are calling a "[race to the bottom](<https://www.thedrum.com/opinion/2024/05/24/ai-already-race-the-bottom>)." Desperate to outmuscle the competition, it seems big tech firms would rather [dazzle investors](<https://www.nytimes.com/2024/05/15/opinion/artificial-intelligence-ai-openai-chatgpt-overrated-hype.html>) with hype than ask "is AI the [right tool](<https://hbr.org/2024/12/is-ai-the-right-tool-to-solve-that-problem>) to solve that problem?"

If we want any hope of keeping the cheating to board games, it's critical that AI developers work with safety, not speed, as their top priority.

**More on AI development:** [*The "First AI Software Engineer" Is Bungling the Vast Majority of Tasks It's Asked to Do*](<https://futurism.com/first-ai-software-engineer-devin-bungling-tasks>)

## Author
At Futurism, I focus on the intersection of technology and power — examining the economics, history, and politics behind today’s dystopian headlines. As a writer, I’m interested in topics ranging from AI’s impact on labor to startups nobody asked for. My prior work includes bylines in Jacobin, Verso, and Blue Labyrinths. My work for Futurism has been cited by publications including Forbes, The Guardian, MIT Technology Review, Time, The Nation, Mother Jones, The Verge, and Wired. I grew up in Michigan, attending Central Michigan University as well as Ball State University, where I earned a master of music. I now live in Brooklyn with my girlfriend and our cat Ziti. On weekends, you can find me hunched over a cold pint arguing geopolitics with the other transplants.

### Author social links  
[Bluesky](<https://bsky.app/profile/joeonhere.bsky.social>)