---
title: "OpenAI’s AI Agent Has to Be Monitored Nonstop to Catch Its Constant Screwups"
description: "OpenAI's new AI agent, Operator, isn't trustworthy or reliable enough to work without constant user supervision."
date: "2025-02-01"
modified: "2025-02-01"
authors:
  - name: "Frank Landymore"
    job_title: "Contributing Writer"
    link: "https://futurism.com/authors/flandymore"
url: "https://futurism.com/the-byte/openai-ai-agent-monitored-nonstop"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "ai agents"
  - "OpenAI"
  - "operator"
  - "the digest"
---

# OpenAI’s AI Agent Has to Be Monitored Nonstop to Catch Its Constant Screwups

![OpenAI's new AI agent, Operator, isn't trustworthy or reliable enough to work without constant user supervision.](<https://futurism.com/wp-content/uploads/2025/01/openai-ai-agent-monitored-nonstop.jpg>)
*\<em\>Image: Updated slightly -- Getty / Futurism\</em\>*

## Baby Brain

OpenAI has finally debuted "Operator," its very own AI agent, a type of autonomous model designed to do digital tasks on your behalf like shopping for groceries online.

But calling it an AI "toddler" might be more fitting. As a *Bloomberg* reporter [describes her experience](<https://www.bloomberg.com/news/newsletters/2025-01-30/openai-s-new-ai-agent-requires-lots-of-adult-supervision>) using OpenAI's new toy, the experimental tech needs a "lot of adult supervision," frequently screwing up and asking for help when it gets stuck.

It's also pretty sluggish, as [echoed by other users](<https://www.reddit.com/r/ChatGPTPro/comments/1i8jln3/i_am_among_the_first_people_to_gain_access_to/>), and slow on the uptake — as a still-developing brain might be.

"For several agonizing moments, I watched as OpenAI's artificially intelligent agent slowly navigated the internet like someone who's had the web described to them in great detail but never actually used it," wrote *Bloomberg*'s Rachel Metz. "I had to monitor it the entire time."

## ParentGPT

Experiences like Metz' suggest there's still a long way to realizing OpenAI's vision — and the industry's at large — of releasing agentic AI models that can serve as virtual employees or personal assistants, supercharging your productivity by doing the work for you.

Typical large language models are limited to words. But AI agents are capable of interacting with their environment, like a user's desktop computer. That potentially enables them to do anything from browse the web — itself a fount of infinite possibilities — to using installed software. In its [announcement](<https://openai.com/index/introducing-operator/>), OpenAI highlighted Operator's usefulness in making reservations, booking flights, and creating shopping lists. The AI model is only available to subscribers of the ChatGPT Pro plan, which costs $200 per month.

If it could do any of those things without help quickly and reliably, it might be a huge time saver. The tech is in a very early stage, however, and [isn't as hands-off as one might hope](<https://futurism.com/openai-asks-permission-important>).

As OpenAI warned, Operator has to ask you for confirmation before actually completing any important tasks — betraying that the tech isn't yet trustworthy enough to be left alone.

## Helicopter Parent

*Bloomberg's* Metz says that Operator successfully handled tasks like getting ice cream delivered, although it required some "guidance and permission," like providing payment info and approving the purchase.

With more serious applications like creating a spreadsheet for a schedule, it frequently messed up the details, she said. OpenAI admitted that Operator still struggles with "complex interfaces" like creating slideshows and managing calendars.

If it can Instacart you some food and do some shopping, cool. But is it worth shelling out $200 a month for?

"Given it was just a test, I was ready and willing to keep a close eye on the product," Metz concluded. "But if OpenAI and its peers want agents to take off, they'll need to convince people that they can trust these services to act autonomously on their behalf. Otherwise, we may decide if we want a job done right, we should just do it ourselves."

**More on OpenAI:** *[OpenAI Hit With Wave of Mockery for Crying That Someone Stole Its Work Without Permission to Build a Competing Product](<https://futurism.com/openai-mockery-stole-work-deepseek>)*

## Author
At Futurism, my work has often centered on bringing a sense of clarity and insight to complex topics ranging from the regulation of emerging technologies to the esoteric ideologies of Silicon Valley executives, while striving not to lose the poetic sense of awe inspired by often-obscure fields like astrophysics and quantum computing. I broke the story of CNET using AI to produce articles that turned out to be riddled with factual errors and plagiarism — a dam-breaking inflection point, as I've reported, that's inspired copycats and endless discourse while beguiling stakeholders ranging from tech giants to purveyors of spam around the web. My work at Futurism has been cited by publications including CBS News, the Los Angeles Times, Vice, Gizmodo, Engadget, the Verge, and Vanity Fair. I grew up in locales ranging from India to China, and now live in the exotic suburbs of Virginia. In my free time, I'm an avid reader of weird sci-fi literature, an aficionado of East Asian cinema, and, regrettably, a relapsed gamer. Allegedly, I’m working on a debut novel, currently untitled.

### Author social links  
[Bluesky](<https://bsky.app/profile/f-w-l.bsky.social>)