---
title: "AI Seems to Do Better on Tasks When Asked to Reflect on Its Mistakes"
description: "A team of AI scientists is claiming that just like humans can solve problems through self-reflection, Large Language Model (LLM) AIs might be able to, too."
date: "2023-03-28"
modified: "2023-03-28"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/ai-reflect-mistakes"
categories:
  - "Artificial Intelligence"
tags:
  - "artificial intelligence"
  - "gpt-4"
  - "large language models"
---

# AI Seems to Do Better on Tasks When Asked to Reflect on Its Mistakes

![A team of AI scientists is claiming that just like humans can solve problems through self-reflection, Large Language Model (LLM) AIs might be able to, too.](<https://futurism.com/wp-content/uploads/2023/03/ai-reflect-mistakes.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

In a not-yet-peer-reviewed [paper](<https://arxiv.org/abs/2303.11366>), a team of researchers from Northeastern University and the Massachusetts Institute of Technology suggests that large language models (LLM) might be able to learn from their own mistakes — just like humans.

Teaching them to do so, they say, might be able to push AI technologies into a new phase of autonomous problem-solving.

"Self-reflection allows humans to efficiently solve novel problems through a process of trial and error," the researchers write in the paper. "Building on recent research, we propose Reflexion, an approach that endows an agent with dynamic memory and self-reflection capabilities to enhance its existing reasoning trace and task-specific action choice abilities."

In other words, their methodology dubbed "Reflexion" is a framework for teaching AI models via prompts to apply a trial-and-error technique to their outputs.

So, just like us, if [at first, they don't succeed](<https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1767250/#:~:text=%E2%80%9CIf%20at%20first%20you%20don,would%20agree%20with%20its%20sentiment.>), they can try, try again.

Testing their new framework was a relatively simple process. The machine, or "agent," was presented with problem-solving tasks and asked to complete them; when it messed up, it was prompted with the Reflexion technique to find those mistakes for *itself* — a process that they claim helps the program evolve, just like humans.

"To achieve full automation, we introduce a straightforward yet effective heuristic that enables the agent to pinpoint hallucination instances, avoid repetition in action sequences, and, in some environments, construct an internal memory map of the given environment," the researchers write in their paper.

Using a series of standardized "decision-making tasks," the researchers found that their methodology was able to greatly improve on a model's given success rates.

The scientists note that their research was conducted using GPT-3 and GPT-3.5-powered AIs — an important consideration, given that OpenAI just released the [much more powerful GPT-4](<https://futurism.com/gpt-4-sparks-of-agi>). Although, in an [accompanying blog post](<https://nanothoughts.substack.com/p/reflecting-on-reflexion>) that discusses the paper, the scientists say that when applied to GPT-4, a "slightly-improvised Reflexion-based GPT-4 agent" was correct 88 percent of the time, outperforming its 67 percent success rate pre-Reflexion.

Again, this paper isn't peer-reviewed, so definitely take the researchers' results with the usual grain of salt.

That said, AI programs mess up a lot, and as they continue to be embedded into workflows across industries and platforms, frameworks for addressing their pitfalls are certainly needed. While this research is more or less an exercise in prompt engineering — rather than addressing the problem of hallucinations from the inside out — it could help in the development of tools that can verify the [infamously unreliable output](<https://futurism.com/the-byte/researchers-gpt-4-accuracy>) of AI language models.

Besides, a little self-reflection never hurt anyone — human or machine.

**More on machine hallucinations:** [*Researchers Find Gpt-4 Is Significantly Less Accurate than Gpt-3*](<https://futurism.com/the-byte/researchers-gpt-4-accuracy>)

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)