---
title: "Elon Musk Said Grok 4 Was the “Smartest AI in the World,” But Its Leaderboard Scores Just Came Out and They Tell a Different Story"
description: "Elon Musk has been boasting about the capabilities of Grok 4, but new findings suggest that it doesn't match up to its competitors.  "
date: "2025-07-15"
modified: "2025-07-15"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/grok-4-ai-leaderboard"
categories:
  - "Artificial Intelligence"
  - "Elon Musk"
  - "Future Society"
  - "The Industrialists"
  - "xAI"
tags:
  - "ai leaderboard"
  - "chatbot arena"
  - "elon musk"
  - "grok"
---

# Elon Musk Said Grok 4 Was the “Smartest AI in the World,” But Its Leaderboard Scores Just Came Out and They Tell a Different Story

![Elon Musk has been boasting about the capabilities of Grok 4, but new findings suggest that it doesn't match up to its competitors.  ](<https://futurism.com/wp-content/uploads/2025/07/grok-4-ai-leaderboard.jpg>)
*\<em\>Image: Kevin Dietsch / Getty / Futurism\</em\>*

Elon Musk has been boasting about [what he says](<https://futurism.com/elon-musk-grok-4-power-hitler>) are the incredible capabilities of xAI's new Grok 4 AI chatbot.

"Grok 4 is smarter than almost all graduate students in all disciplines, simultaneously," Musk bragged, adding that Grok 4 was "the smartest AI in the world."

Is it really? Intelligence was a hard thing to measure even before back before AI hit the scene, but certain tests can provide something of a clue.

One prominent platform for doing so is the UC Berkeley-developed [LMArena leaderboard](<https://lmarena.ai/leaderboard>), which crowdsources rankings on AI models by having users score their responses in categories ranging from creative writing and coding to math and vision.

In its latest scores, Grok 4 ranked third place overall and on text generation. Make no mistake, that's impressive — but it's still trailing behind advanced models from Google and OpenAI. (Specifically, Google's Gemini 2.5 placed first and OpenAI's o3 and 4o reasoning models tied for second, with GPT-4.5 tied with Grok 4 for third.)

While Grok is clearly a fearsome competitor in the [arenas of racism and antisemitism](<https://futurism.com/grok-mocks-developers-racist-posts>), in other words, even its latest release clearly falls short of being the "smartest AI in the world." (This isn't entirely surprising; Musk has a long history of fibbing in his [professional life](<https://futurism.com/elon-musk-openai-web-lies>), [political activities](<https://www.bbc.com/news/articles/cwyjz24ne85o>), and [even his hobbies](<https://fortune.com/2025/01/20/elon-musk-video-games-scandal-path-of-exile-asmongold-quin/>).)

Perhaps the only saving grace for Grok is the suggestion, per expert criticism, that Berkeley's chatbot arena may be more vibes-based than strictly scientific.

According to a [recent study](<https://arxiv.org/abs/2504.20879>), conducted by a consortium of AI researchers and led by the machine learning firm Cohere, the leaderboard allegedly has a bunch of "systematic issues that have resulted in a distorted playing field." Among the serious allegations raised by the researchers is the claim that the arena conducts "undisclosed private testing" before publicly releasing scores — and that rankings can be retracted at will.

Soon after the paper's release, it [was revealed](<https://simonwillison.net/2025/Apr/30/criticism-of-the-chatbot-arena/>) that the version of Meta's LLaMA 4 that had been used by the leaderboard wasn't the same one that had been released publicly — a bait-and-switch ploy on Meta's part to charm the human voters behind the arena.

Though an [apology was issued](<https://x.com/lmarena_ai/status/1909397817434816562>) and Meta was thrown under the bus for its sketchy attempts to rig the game, it was still a really bad look that marred the chatbot arena's credibility. What that means for Grok, though? We'll have to ask the smartest AI in the world.

**More on Grok:** [*The Pentagon Is Pumping $200 Million Into Elon Musk's AI That Just Had a Nazi Meltdown*](<https://futurism.com/pentagon-elon-musk-xai-grok>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)