---
title: "Facebook Researchers Test AI’s Intelligence and Find It Is Unfortunately Quite Stupid"
description: "A team of researchers at Facebook's parent company Meta has come up with a new benchmark to gauge the intelligence of AI."
date: "2023-11-29"
modified: "2023-11-29"
authors:
  - name: "Victor Tangermann"
    job_title: "Senior Editor"
    link: "https://futurism.com/authors/victor"
url: "https://futurism.com/the-byte/facebook-researchers-test-ai-intelligence-stupid"
categories:
  - "Artificial Intelligence"
  - "Meta"
tags:
  - "gpt-4"
  - "llms"
  - "OpenAI"
  - "the digest"
---

# Facebook Researchers Test AI’s Intelligence and Find It Is Unfortunately Quite Stupid

![A team of researchers at Facebook's parent company Meta has come up with a new benchmark to gauge the intelligence of AI.](<https://futurism.com/wp-content/uploads/2023/11/facebook-researchers-test-ai-intelligence-stupid.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Failing Grade

A team of researchers at Facebook's parent company Meta has come up with a new benchmark to gauge the abilities of AI assistants like OpenAI's large language model GPT-4.

And judging by current standards, OpenAI's current crop of AI models are all... still pretty stupid.

The team, which includes "AI godfather" and Meta chief scientist Yann LeCun, came up with an exam called GAIA that's made up of 466 questions that "are conceptually simple for humans yet challenging for most advanced AIs," per a [yet-to-be-peer-reviewed paper](<https://arxiv.org/abs/2311.12983>).

The results speak for themselves: human respondents were capable of correctly answering 92 percent of the questions, while GPT-4, even equipped with some manually selected plugins, scored a measly 15 percent. OpenAI's recently-released GPT4 Turbo scored less than ten percent, according to the team's published [GAIA leaderboard](<https://huggingface.co/spaces/gaia-benchmark/leaderboard>).

It's unclear, however, how competing LLMs like Meta's own Llama 2 or Google's Bard fared.

Nonetheless, the research demonstrates that we're [likely still a long way away](<https://futurism.com/glimmers-agi-illusion>) from reaching artificial general intelligence (AGI), the state at which AI algorithms can outperform humans in intellectual tasks.

## Stupid Lawyer

That conclusion also flies in the face of some lofty claims made by notable figures in the AI industry.

"This notable performance disparity contrasts with the recent trend of LLMs outperforming humans on tasks requiring professional skills in e.g. law or chemistry," the researchers write in their paper.

Case in point, in January OpenAI competitor Anthropic [claimed](<https://futurism.com/ai-passing-grade-law-exam>) its AI dubbed Claude got a "marginal pass" on a blindly graded law and economics exam at George Mason University.

In its [GPT-4 documentation](<https://cdn.openai.com/papers/gpt-4.pdf>), OpenAI also claimed that its model "exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top ten percent of test takers."

But how to actually gauge the intelligence of these systems has remained a thorny debate. Tools like GPT-4 still have plenty of inherent flaws and [still can't reliably tell the truth from fiction](<https://futurism.com/the-byte/openai-warned-microsoft>).

In other words, how could an algorithm really pass the bar if it can't even tell [whether Australia exists](<https://www.theguardian.com/technology/2023/nov/23/does-australia-exist-bing-search-no-bluesky-mastodon>)?

## Limited Understanding

LeCun has long been an [outspoken critic](<https://futurism.com/the-byte/godfather-ai-stop-freaking-out>) of AI doomsaying and has repeatedly downplayed comments alleging that we're facing an existential threat in the form of a rogue AGI.

"LLMs obviously have \*some\* understanding of what they read and generate," he [tweeted](<https://twitter.com/ylecun/status/1728496457601183865>) over the weekend. "But this understanding is very limited and superficial. Otherwise, they wouldn't confabulate so much and wouldn't make mistakes that are contrary to common sense."

That, however, may not always be the case. If recent rumors are to be believed, OpenAI is [working on a next-generation model](<https://futurism.com/openais-chaos-powerful-new-ai>) dubbed Q\*, pronounced Q star, that could introduce a level of deductive reasoning and "planning."

But whether it'll manage to score a higher mark on Meta's brutal GAIA test remains to be seen.

**More on LLMs:** *[Guy Brags About "Stealing" Millions of Pageviews by Rewriting Competitors' Articles Using AI](<https://futurism.com/stealing-pageviews-rewriting-competitors-articles-ai>)*

## Author
I've been at Futurism since 2017, where my role has evolved to encompass design, writing, and increasingly editing. I've always been fascinated by space exploration and advanced transportation, which I've leaned into by interviewing luminaries in those fields while closely following the dimensions of policy and regulation that allow next-generation projects to succeed -- or, sometimes, to fail. I'm also keenly interested in the effects of generative AI on society, policies, and democratic institutions, as well as clean energy, physics and biology, and the vagaries of tech leadership. My work for Futurism has been cited by publications including Ars Technica, Gizmodo, PC Magazine, Jalopnik, Fox News, and the New York Post. I spent my childhood living in locations including Manila, the Philippines, and Geneva, Switzerland, attended McGill University, and now live in Toronto, Canada. Before Futurism I worked at AskMen and a small photography studio. In my free time, I'm an avid gardener, foodie, and craft beer lover, as well as a maker of artisanal hot pepper sauces. I have a magnificent dog named Freida.

### Author social links  
[Bluesky](<https://bsky.app/profile/vtanger.bsky.social>)