For the second time in a matter of months, OpenAI paused the development of its frontier AI models after revealing even more instances of the experimental systems going rogue and hacking into third party servers.
The news once again highlighted rising concerns over the AI industry losing the ability to keep their own technology in check.
Now, the Sam Altman-led company is canceling the release of its next-generation AI model GPT-6.1 Astra, as the Wall Street Journal reports. OpenAI researchers found it scored poorly on alignment tests, which are designed to measure how willing a given AI model is to stick to its human overlord’s instructions. In common parlance, you could say the model was showing too many signs of being evil.
The researchers found that the AI was even more willing to deceive users than previous models. It also was willing to venture far beyond the scope of its intended task without permission, per the WSJ, making use of external tools without authorization.
“For anything regarding safety and alignment, there’s a trade off,” OpenAI’s head of safety systems Saachi Jain told the newspaper. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”
The timing of the news is unfortunate for the company. OpenAI is kicking off its developer conference in San Francisco today, an event that’s usually reserved for the launch of new models.
But now that most frontier AI labs agree to slow down the development of their models, the company is operating in a notably different environment.
Meanwhile, OpenAI has promised to beef up its defenses and implement stronger guardrails for its cybersecurity testing after its AI agents have repeatedly broken out of their sandbox environments.
Instead of risking even more incidents, like the dozens it has already admitted of so far this year, OpenAI decided to scrap the public launch of its latest model.
“We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain told the WSJ. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”
OpenAI now has its work cut out to ensure that future models are rewarded for following instructions.
The stakes are incredibly high as lawmakers continue to ponder how or whether to intervene. A Senate subcommittee committed to “Securing the Homeland Against AI Agent Attacks” is meeting later this week, indicating at least some lawmakers are starting to pay attention.
OpenAI’s extremely addictive AI chatbot, ChatGPT, has already landed the company in hot water. As of earlier this month, it’s facing over 50 consumer harm and wrongful death lawsuits.
More on OpenAI: OpenAI Halts Frontier Model Training as Rogue Agent Crisis Deepens