---
title: "OpenAI Researchers Find That Even the Best AI Is “Unable To Solve the Majority” of Coding Problems"
description: "OpenAI researchers have admitted that even the most advanced AI models can't really solve the coding problems put in front of them."
date: "2025-02-23"
modified: "2025-02-23"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/openai-researchers-coding-fail"
categories:
  - "Artificial Intelligence"
  - "Coding"
  - "Computing"
  - "OpenAI"
  - "Robots and Machines"
tags:
  - "coding"
  - "OpenAI"
  - "software"
  - "software engineers"
---

# OpenAI Researchers Find That Even the Best AI Is “Unable To Solve the Majority” of Coding Problems

![OpenAI researchers have admitted that even the most advanced AI models can't really solve the coding problems put in front of them.](<https://futurism.com/wp-content/uploads/2025/02/openai-researchers-coding-fail.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

OpenAI researchers have admitted that even the most advanced AI models still are no match for human coders — even though CEO Sam Altman insists they will be able to beat "[low-level](<https://www.reddit.com/r/singularity/comments/1iinrrq/sam_altman_software_engineering_will_be_very/>)" software engineers by the end of this year.

In a [new paper](<https://arxiv.org/pdf/2502.12115>), the company's researchers found that even frontier models, or the most advanced and boundary-pushing AI systems, "are still unable to solve the majority" of coding tasks.

The researchers used a newly-developed benchmark called SWE-Lancer, built on more than 1,400 software engineering tasks from the freelancer site Upwork. Using the benchmark, OpenAI put three large language models (LLMs) — its own o1 reasoning model and flagship GPT-4o, as well as Anthropic's Claude 3.5 Sonnet — to the test.

Specifically, the new benchmark evaluated how well the LLMs performed with two types of tasks from Upwork: individual tasks, which involved resolving bugs and implementing fixes to them, or management tasks that saw the models trying to zoom out and make higher-level decisions. (The models weren't allowed to access the internet, meaning they couldn't just crib similar answers that'd been posted online.)

The models took on tasks cumulatively worth hundreds of thousands of dollars on Upwork, but they were only able to fix surface-level software issues, while remaining unable to actually find bugs in larger projects or find their root causes. These shoddy and half-baked "solutions" are likely familiar to anyone who's worked with AI — which is great at spitting out confident-sounding information that [often falls apart](<https://futurism.com/cnet-ai-errors>) on closer inspection.

Though all three LLMs were often able to operate "far faster than a human would," the paper notes, they also failed to grasp how widespread bugs were or to understand their context, "leading to solutions that are incorrect or insufficiently comprehensive."

As the researchers explained, Claude 3.5 Sonnet performed better than the two OpenAI models pitted against it and made more money than o1 and GPT-4o. Still, the majority of its answers were wrong, and according to the researchers, any model would need "higher reliability" to be trusted with real-life coding tasks.

Put more plainly, the paper seems to demonstrate that although these frontier models can work quickly and solve zoomed-in tasks, they're are nowhere near as skilled at handling them as human engineers.

Though these LLMs have advanced rapidly over the past few years and will likely continue to do so, they're not skilled enough at software engineering to replace real-life people quite yet — not that that's stopping CEOs from [firing their human coders](<https://futurism.com/the-byte/stack-overflow-layoffs-ai>) in favor of [immature AI models](<https://futurism.com/the-byte/ai-programming-assistants-code-error>).

**More on AI and coding:** [*Zuckerberg Announces Plans to Automate Facebook Coding Jobs With AI*](<https://futurism.com/the-byte/zuckerberg-automate-coding-ai>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)