---
title: "Scientists Preparing “Humanity’s Last Exam” to Test Powerful AI"
description: "AI experts are calling for submissions for the \"hardest and broadest set of questions ever\" to try to stump AI systems."
date: "2024-09-18"
modified: "2024-09-18"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/humanitys-last-exam-ai-benchmarks"
categories:
  - "Artificial Intelligence"
tags:
  - "ai benchmarks"
  - "ai safety"
  - "the digest"
  - "training data"
---

# Scientists Preparing “Humanity’s Last Exam” to Test Powerful AI

![AI experts are calling for submissions for the "hardest and broadest set of questions ever" to try to stump AI systems.](<https://futurism.com/wp-content/uploads/2024/09/humanitys-last-exam-ai-benchmarks.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Final Exam

AI experts are calling for submissions for the "hardest and broadest set of questions ever" to try to stump today's most advanced artificial intelligence systems — as well as those that are still coming.

As [*Reuters* reports](<https://www.reuters.com/technology/artificial-intelligence/ai-experts-ready-humanitys-last-exam-stump-powerful-tech-2024-09-16/>), this test — known in the field, memorably, as "Humanity's Last Exam" — is being crowdsourced by the Center for AI Safety (CAIS) and the training data labeling firm Scale AI, which over the summer [raised a cool billion dollars](<https://fortune.com/2024/05/21/scale-ai-funding-valuation-ceo-alexandr-wang-profitability/>) for an overall value of $14 billion.

*Reuters* points out that submissions for this "exam" were opened just a day after results from [OpenAI's new o1 model preview](<https://futurism.com/openai-strawberry-o1-mistakes>) dropped. As CAIS executive director Dan Hendryks notes, o1 seems to have "destroyed the most popular reasoning benchmarks."

Back in 2021, [Hendrycks co-authored](<https://arxiv.org/abs/2105.09938>) [two papers](<https://openreview.net/forum?id=dNy_RKzJacY>) with AI testing proposals that would evaluate whether models could out-quiz undergraduates. At the time, the AI systems being tested were spouting off answers nearly at random, but as Hendrycks notes, the models of today have "crushed" the 2021 tests.

## Abstract Thinking

While the 2021 testing criteria primarily grilled the AI systems on math and social studies, "Humanity's Last Exam" will, as the CAIS executive director said, incorporate abstract reasoning to make it harder. The two institutions organizing the test are also planning to keep the test criteria confidential and not opening it up to the public, to make sure the answers don't end up in any AI training data.

Due November 1, experts in fields as far-flung as rocketry and philosophy are being encouraged to submit questions that would be difficult for those outside their areas of expertise to answer. After undergoing peer review, winners will be offered co-authorship of a paper associated with the test and prizes up to $5,000 sponsored by Scale AI.

While the organizers are casting a very wide net for the types of questions they're seeking, they told *Reuters* that there's one thing that will not be on the exam: anything about weapons, because it's too dangerous for AI to know about.

**More on advanced AI:** [*OpenAI's Strawberry "Thought Process" Sometimes Shows It Scheming to Trick Users*](<https://futurism.com/openai-strawberry-thought-process-scheming>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)