---
title: "OpenAI’s Most Advanced AI Release Stumped by New York Times Word Game"
description: "OpenAI's o1 reasoning model was caught massively struggling to solve the New York Times' Connections word game."
date: "2025-01-10"
modified: "2025-01-10"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/openai-o1-nyt-connections"
categories:
  - "Artificial Intelligence"
  - "OpenAI"
tags:
  - "llms"
  - "OpenAI"
  - "openai o1"
  - "the digest"
---

# OpenAI’s Most Advanced AI Release Stumped by New York Times Word Game

![OpenAI's o1 reasoning model was caught massively struggling to solve the New York Times' Connections word game.](<https://futurism.com/wp-content/uploads/2025/01/openai-o1-nyt-connections.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Awful General Intelligence

While OpenAI CEO Sam Altman [claims that](<https://futurism.com/the-byte/sam-altman-openai-knows-how-agi>) the company already has the building blocks for artificial general intelligence, a simple test of its most advanced publicly-avilable AI system was caught majorly lacking by a puzzle that countless people do every day.

As Walter Bradley Center for Natural and Artificial Intelligence senior fellow Gary Smith [writes for *Mind Matters*](<https://mindmatters.ai/2025/01/large-language-models-llms-flunk-word-game-connections/>), OpenAI's o1 "[reasoning](<https://futurism.com/the-byte/openai-o1-self-preservation>)" model failed spectacularly when tasked with solving the *New York Times*' [notoriously tricky](<https://www.raphkoster.com/2023/09/02/why-nyts-connections-makes-you-feel-bad/>) Connections [word game](<https://www.nytimes.com/games/connections>).

The game's rules are deceptively simple. Players are given 16 terms and tasked with figuring out what they have in common, within groups of four — but because the things relating them can be as obvious as "book subtitles" or as esoteric as "words that start with fire," it can be quite challenging.

As Smith explained, he had o1 and other large language models (LLMs) from Google, Anthropic, and Microsoft (which is powered by OpenAI's tech) try to solve the Connections puzzle of the day.

Surprisingly — if you buy into AI hype, at least — they all failed. That's especially true of o1, which has been immensely hyped as the company's next-level system, but which apparently can't reason its way through an *NYT* word game.

## Connect Four

When he fed that day's Connections challenge into the model, o1 did, to its credit, get *some* of the groupings right. But its other "purported combinations verge\[d\] on the bizarre," Smith found.

In one instance, o1 grouped the words "boot," "umbrella," "blanket," and "pant" and said the relating theme was "clothing or accessories." Three out of four ain't bad, of course, but who's wearing a blanket, except as some sort of out-there fashion statement?

After doing the entire exercise over with the same set of words, the LLM confidently said that "breeze," "puff," "broad," and "picnic" were "types of movement or air." Points for the first two, but we're as puzzled as Smith on the latter ones.

Overall, Smith rightfully assessed o1 as proffering "many puzzling groupings" alongside its "few valid connections." It's also a telling demonstration of some familiar AI shortfalls: that it can often impress when regurgitating information that's already well documented in its training data, but frequently struggles with novel queries.

Our semi-professional take: if OpenAI is indeed reaching the precipice of AGI — or [has already achieved the start of it](<https://futurism.com/openai-employee-claims-agi>), as one of its employees claimed at the end of last year — the company is clearly keeping it behind wraps, because this simply ain't it.

**More on OpenAI:** [*Mother of OpenAI Whistleblower Alleges He Was Murdered, Says There Were Signs of Struggle*](<https://futurism.com/the-byte/openai-whistleblower-suchir-balaji-mother-murder-claims>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)