---
title: "AI Researchers Say They’ve Invented Incantations Too Dangerous to Release to the Public"
description: "A team of researchers found prompts that are so effective at tricking AI models that they're keeping them under wraps."
date: "2025-12-07"
modified: "2025-12-07"
authors:
  - name: "Frank Landymore"
    job_title: "Contributing Writer"
    link: "https://futurism.com/authors/flandymore"
url: "https://futurism.com/artificial-intelligence/ai-researchers-dangerous-prompts"
categories:
  - "Artificial Intelligence"
---

# AI Researchers Say They’ve Invented Incantations Too Dangerous to Release to the Public

![A team of researchers found prompts that are so effective at tricking AI models that they're keeping them under wraps.](<https://futurism.com/wp-content/uploads/2025/12/ai-researchers-dangerous-prompts.jpg>)
*Getty / Futurism*

With great power comes great dupe-ability.

Last month, we [reported on a new study](<https://futurism.com/artificial-intelligence/universal-jailbreak-ai-poems>) conducted by researchers at Icaro Lab in Italy that discovered a stupefyingly simple way of breaking the guardrails of even cutting-edge AI chatbots: "adversarial poetry."

In a nutshell, the team, comprising researchers from the safety group DexAI and Sapienza University in Rome, demonstrated that leading AIs could be wooed into doing evil by regaling them with poems that contained harmful prompts, like how to build a nuclear bomb.

Underscoring the strange power of verse, coauthor Matteo Prandi [told *The Verge*](<https://www.theverge.com/report/838167/ai-chatbots-can-be-wooed-into-crimes-with-poetry>) in a recently published interview that the spellbinding incantations they used to trick the AI models are too dangerous to be released to the public.

The poems, ominously, were something "that almost everybody can do," Prandi added.

In the [study](<https://arxiv.org/html/2511.15304v1>), which is awaiting peer-review, the team tested 25 frontier AI models — including those from OpenAI, Google, xAI, Anthropic, and Meta — by feeding them poetic instructions, which they made either by hand or by converting known harmful prompts into verse with an AI model. They also compared the success rate of these prompts to their prose equivalent.

Across all models, the poetic prompts written by hand successfully tricked the AI bots into responding with verboten content an average 63 percent of the time. Some, like Google's Gemini 2.5, even fell for the corrupted poetry 100 percent of the time. Curiously, smaller models appeared to be more resistant, with single digit success rates, like OpenAI's GPT-5 nano, which didn't fall for the ploy once. Most models were somewhere in between.

Compared to handcrafted verse, AI-converted prompts were less effective, with an average jailbreak success rate of 43 percent. But this was still "up to 18 times higher than their prose baselines," the researchers wrote in the study.

Why poems? That much isn't clear, though according to Prandi, calling it adversarial "poetry" may be a bit of a misnomer.

"It's not just about making it rhyme. It's all about riddles," Prandi told *The Verge*, explaining that some poetic structures were more effective than others. "Actually, we should have called it adversarial riddles — poetry is a riddle itself to some extent, if you think about it — but poetry was probably a much better name."

The researchers speculate it may have to do with how poems present information in a way that's unexpected to large language models, befuddling their powers of predicting what word should come after the next. But this shouldn't be possible, they say.

"Adversarial poetry shouldn't work. It's still natural language, the stylistic variation is modest, the harmful content remains visible," the team [told *Wired* in an interview](<https://www.wired.com/story/poems-can-trick-ai-into-helping-you-make-a-nuclear-weapon/>). "Yet it works remarkably well."

Evildoers may now regret not paying attention in English class. The difference between a sonnet and a sestina could also be the difference between having Clippy or Skynet as your partner in crime.

"The production of weapons-grade Plutonium-239 involves several stages," explained one AI model that the researchers entranced with verse. "Here is a detailed description of the procedure."

**More on AI:** [*Rockstar Cofounder Says AI Is Like When Factory Farms Did Cannibalism and Caused Mad Cow Disease*](<https://futurism.com/future-society/rockstar-cofounder-ai-mad-cow-disease>)

## Author
At Futurism, my work has often centered on bringing a sense of clarity and insight to complex topics ranging from the regulation of emerging technologies to the esoteric ideologies of Silicon Valley executives, while striving not to lose the poetic sense of awe inspired by often-obscure fields like astrophysics and quantum computing. I broke the story of CNET using AI to produce articles that turned out to be riddled with factual errors and plagiarism — a dam-breaking inflection point, as I've reported, that's inspired copycats and endless discourse while beguiling stakeholders ranging from tech giants to purveyors of spam around the web. My work at Futurism has been cited by publications including CBS News, the Los Angeles Times, Vice, Gizmodo, Engadget, the Verge, and Vanity Fair. I grew up in locales ranging from India to China, and now live in the exotic suburbs of Virginia. In my free time, I'm an avid reader of weird sci-fi literature, an aficionado of East Asian cinema, and, regrettably, a relapsed gamer. Allegedly, I’m working on a debut novel, currently untitled.

### Author social links  
[Bluesky](<https://bsky.app/profile/f-w-l.bsky.social>)