---
title: "What Actually Happens When Programmers Use AI Is Hilarious, According to a New Study"
description: "As AI takes the programming world by storm, experienced developers are still the gold standard — and they're probably better off without algorithmic assistance, too. As Ars Technica flags, a new study from the Berkeley-based nonprofit Model Evaluation and Threat Research (METR) found that programmers were actually slower when using AI assistance tools than they were just doing tasks straight. In the study, 16 programmers were given roughly 250 coding tasks and asked to either use no AI assistance or use  what METR called \"early-2025 AI tools\" like Anthropic's Claude and Cursor Pro, which both are regularly used for such applications. Despite […]"
date: "2025-07-16"
modified: "2025-07-16"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/ai-coding-programmers-reality"
categories:
  - "Artificial Intelligence"
tags:
  - "ai benchmarks"
  - "ai coding"
  - "Programming"
  - "vibe coding"
---

# What Actually Happens When Programmers Use AI Is Hilarious, According to a New Study

![As AI takes the programming world by storm, experienced devs are still the gold standard — and they're better off without AI assistance, too.](<https://futurism.com/wp-content/uploads/2025/07/ai-coding-programmers-reality.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

AI has taken the programming world by storm, with a [flurry of speculation](<https://www.zdnet.com/article/will-ai-replace-software-engineers-it-depends-on-who-you-ask/>) about the tech replacing human coders, and Google's CEO [recently claiming](<https://futurism.com/the-byte/google-ceo-code-ai>) that 25 percent of the company's code is now AI-generated.

But it's possible that in practice, AI is actually hindering efficient software development. As [flagged by *Ars Technica*](<https://arstechnica.com/ai/2025/07/study-finds-ai-tools-made-open-source-software-developers-19-percent-slower/>), a [new study](<https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/>) from the nonprofit Model Evaluation and Threat Research (METR) found that in practice, programmers are actually *slower* when using AI assistance tools than making do without them.

In the study, 16 programmers were given roughly 250 coding tasks and asked to either use no AI assistance, or employ what METR characterized as "early-2025 AI tools" like Anthropic's Claude and Cursor Pro. The results were surprising, and perhaps profound: the programmers actually spent 19 percent *more* time when using AI than when forgoing it.

When measuring the programmers' screen time, the METR team found that when using AI tools, their subjects did indeed spend less time actively coding, debugging, researching, or testing — but that was because they instead spent their time "reviewing AI outputs, prompting AI systems, and waiting for AI generations."

Ultimately, the AI-assisted cohort accepted less than 44 percent of the tips provided by the tools without any modification, and nine percent of the total time they spent on tasks was eaten up by fixing the AI's outputs. (That's not entirely surprising; companies that laid off people to replace them with AI are now having to hire new contractors to fix [the technology's mistakes](<https://futurism.com/companies-fixing-ai-replacement-mistakes>).)

Despite the results, however, the programmers in the study believed initially that AI would reduce nearly a quarter of the time spent on tasks — and afterward, they still thought those tools sped them up by 20 percent.

Perhaps contributing to the disconnect between expectation and reality with AI coding are [all those benchmarks](<https://aider.chat/docs/leaderboards/>) claiming that these tools, and others like OpenAI's o3 reasoning model and Google's Gemini, are spitting out immaculate code at record speeds. As *Ars* notes, however, those benchmarks rely on "synthetic, algorithmically scorable tasks created specifically" for such tests, which may be a poor reflection of the messy world of actual coding.

This isn't the first time that the narrative of AI's dominance in coding has been rattled by research findings. Earlier this year, for instance, [OpenAI researchers released a paper](<https://arxiv.org/pdf/2502.12115>) declaring, based on benchmarking tests from real-world coding tasks, that even the most advanced large language models "[are still unable to solve the majority](<https://futurism.com/openai-researchers-coding-fail>)" of problems.

AI is also giving rise to other unintended consequences in the world of software development. Untrained programmers who engage in so-called "[vibe coding](<https://x.com/karpathy/status/1886192184808149383?lang=en>)," or writing and fixing code by describing what they want to an AI, are not only screwing up their work itself, but also self-sabotaging by introducing [severe cybersecurity risks](<https://futurism.com/problem-vibe-coding>) to the finished product as well.

With so many tech workers being [laid off in favor of automation](<https://www.bloodinthemachine.com/p/how-ai-is-killing-jobs-in-the-tech-f39>), it stands to reason that code generated after such firings is [less accurate and secure](<https://cyberscoop.com/vibe-coding-ai-cybersecurity-llm/>) than it was when humans were writing it — but thus far, that hasn't seemed to matter much to the [people doing the job cuts](<https://www.wsj.com/tech/ai/ai-white-collar-job-loss-b9856259>).

**More on AI coding:** [*Researchers Trained an AI on Flawed Code and It Became a Psychopath*](<https://futurism.com/openai-bad-code-psychopath>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)