There are a litany of reasons to take the output of generative AI with a heavy dose of salt. Its answers are known to sound authoritative, convincing, and elaborate — while simultaneously being completely and utterly wrong.
It’s a lesson astrophysicist and science communicator Paul Sutter learned the hard way, and was brave enough to share with the world. In a article for the publication Nautilus earlier this month, Sutter retold a truly harrowing story that feels like something yanked straight out of a stress-induced nightmare.
In February, he gave a presentation in front of a room full of “collaborators,” revealing a new update for his algorithm designed to spot voids, or empty regions between galaxies.
The new version was “ten times faster than the old one” had a “more sophisticated way of dealing with the ugly realities of an actual data set” and could “handle surveys a hundred times larger,” Sutter recalled.
It was also in part the result of extensive sessions consulting an AI to write the update’s code — “vibe coding,” in the parlance — which he recalled as “indispensable.”
But just ten minutes into his presentation, a collaborator interjected telling that “something seemed off to them.” As it turns out, the way the new algorithm handled the edges of the surveys was “wrong,” Sutter admitted.
“It wasn’t a typo, and it wasn’t a missing citation or a factor of two,” he wrote. “It was subtle, but it was very wrong, and everything downstream of it was also wrong, and I had shared the whole thing in a room full of people who trusted me.”
Sutter’s cringe-inducing tale perfectly illustrates how many users of AI tools are putting themselves at risk without knowing it, including in the most rarefied echelons of academia. The science world has been inundated with poorly-researched and often unedited AI slop, a major reckoning forcing academics to be responsible for everything they publish — including any potentially embarrassing hallucinations.
The AI coding tool Sutter was using “sounded like it understood,” he wrote, despite being “little more than a sophisticated next-word predictor.”
“An LLM’s fluency is not an accident, and it is not an emergent mystery,” Sutter wrote. “It’s a trait we bred, the way we bred wolves into dogs that watch our faces when we open the treat bag.”
“So how do we deploy a tool that is sometimes wrong but always pleasing? How do we trust AI?” he concluded. “The simple answer is: don’t.”
Sutter likened our extensive use of AI tools to alchemy, an ancient practice that was devised by what he calls “pre-scientists.”
“We are now in the pre-chemistry era of AI,” the astrophysicist declared. “The crucible is closed, and like the alchemists we are not going to stop using it.”
To get to the truth, he says, we must carefully trace an AI’s chain of reasoning and audit its every output. Following his embarrassing slip-up in February, Sutter vowed that he works “differently now” and is ready to immediately “distrust” anything an AI says.
It’s a cautionary tale that even some of the most gifted thinkers can easily be tempted by the allure of AI.
As such, when we pasted his latest piece for Nautilus into AI detecting tool Pangram — which is far from perfect — it informed us that 56 percent of the text appeared to be written by an AI.
More on AI hallucinations: Academics in Meltdown Now That They’re Responsible for AI Hallucinations in Their Research Papers