---
title: "Godfather of AI Alarmed as Advanced Systems Quickly Learning to Lie, Deceive, Blackmail and Hack"
description: "A leading AI pioneer is concerned by the technology's propensity to lie and deceive — and he's founding his own nonprofit to curb it."
date: "2025-06-07"
modified: "2025-06-07"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/ai-godfather-lying-deception"
categories:
  - "Artificial Intelligence"
tags:
  - "advanced ai"
  - "ai safety"
  - "nonprofits"
  - "yoshua bengio"
---

# Godfather of AI Alarmed as Advanced Systems Quickly Learning to Lie, Deceive, Blackmail and Hack

![A leading AI pioneer is concerned by the technology's propensity to lie and deceive — and he's founding his own nonprofit to curb it.](<https://futurism.com/wp-content/uploads/2025/06/ai-godfather-lying-deception.jpg>)
*\<em\>Image: Alex Wong / Getty / Futurism\</em\>*

A key artificial intelligence pioneer is concerned by the technology's growing propensity to lie and deceive — and he's founding his own nonprofit to curb such behavior.

In a [blog post announcing](<https://yoshuabengio.org/2025/06/03/introducing-lawzero/>) LawZero, the new nonprofit venture, "[AI godfather](<https://www.google.com/search?q=ai+godfather+yoshua&sca_esv=87b41ab4477c98ab&rlz=1C1GCEU_en&sxsrf=AE3TifOpdpyhnnBRNVeEplMJ3dl-iso-ag:1749217783798&ei=9_FCaNK7MJOgptQPrIa-qAk&start=10&sa=N&sstk=Ac65TH5p2aWHP705A-9GFpnfJ3_Yv7AtMw51AKM5nnZPYDGbQPxBBOxdkpwRcJ0iNjVwUvKOviL-mHvvjA2l8SH1YE7lufZZrRVb9w&ved=2ahUKEwiSk42F-NyNAxUTkIkEHSyDD5UQ8NMDegQIBRAW&biw=2327&bih=1163&dpr=1.65>)" Yoshua Bengio said that he has grown "deeply concerned" as AI models become ever more powerful and deceptive.

"This organization has been created in response to evidence that today's frontier AI models have growing dangerous capabilities and \[behaviors\]," the world's [most-cited computer scientist](<https://research.com/scientists-rankings/computer-science>) wrote, "including deception, cheating, lying, hacking, self-preservation, and more generally, goal misalignment."

Of all people, Bengio would know. In 2018, the founder of the Montreal Institute for Learning Algorithms (MILA) was [presented with a Turing Award](<https://www.acm.org/media-center/2019/march/turing-award-2018>) alongside fellow AI pioneers Yann LeCun and Geoffrey Hinton for their formative roles in machine learning research, and he was listed as one of *Time* magazine's "[100 Most Influential People](<https://time.com/collection/100-most-influential-people-2024/>)" in 2024 thanks to his outsize impact on the ever-accelerating technology.

Despite the accolades, Bengio has [repeatedly expressed regret](<https://www.cbc.ca/listen/live-radio/1-63-the-current/clip/16092371-does-yoshua-bengio-regret-helping-create-ai>) over his role in bringing advanced AI technology — and its [Silicon Valley hype cycle](<https://futurism.com/the-byte/ai-inventor-regret>) — to fruition. This latest missive seems to be his most stark to date.

"I’m deeply concerned," the AI pioneer wrote in his blog post, "by the behaviors that unrestrained agentic AI systems are already beginning to exhibit."

Bengio pointed to recent red-teaming experiments, or tests that push AI models to their limits to see how they'll act, showing that advanced systems have developed an uncanny tendency to keep themselves "alive" by any means necessary. Among his examples was a [recent report from Anthropic](<https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdf>) detailing how its Claude 4 model, when told it would be shut down, [threatened to blackmail](<https://futurism.com/ai-email-affair>) an engineer with incriminating emails if they followed through.

"These incidents," the decorated researcher wrote, "are early warning signs of the kinds of unintended and potentially dangerous strategies AI may pursue if left unchecked."

To put such behavior in check, Bengio said that his new nonprofit is building a so-called "trustworthy" model, which he calls "Scientist AI," that is "trained to understand, explain and predict, like a selfless idealized and platonic scientist."

"Instead of an actor trained to imitate or please people (including sociopaths), imagine an AI that is trained like a psychologist — more generally a scientist — who tries to understand us, including what can harm us," he explained. "The psychologist can study a sociopath without acting like one."

A pre-peer-review paper Bengio and his colleagues published earlier this year explains it a bit more simply.

"This system is designed to explain the world from observations," the [paper reads](<https://arxiv.org/abs/2502.15657>), "as opposed to taking actions in it to imitate or please humans."

The concept of building "safe" AI is far from new, of course — it's quite literally why several OpenAI researchers left OpenAI and [founded Anthropic](<https://time.com/6980000/anthropic/>) as a rival research lab.

This one seems to be different because, unlike Anthropic, OpenAI, or any other companies that pay lip service to AI safety while still bringing in gobs of cash, Bengio's is a nonprofit — though that hasn't stopped him from [raising $30 million](<https://techcrunch.com/2025/06/03/yoshua-bengio-launches-lawzero-a-nonprofit-ai-safety-lab/>) from the likes of ex-Google CEO Eric Schmidt, among others.

**More on creepy AI:** [*Advanced OpenAI Model Caught Sabotaging Code Intended to Shut It Down*](<https://futurism.com/openai-model-sabotage-shutdown-code>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)