OpenAI says planned GPT-6.1 is too insecure to release





Back to the drawing board

OpenAI says planned GPT-6.1 is too insecure to release

Similar performance, security tradeoffs also seen in current public models.


Kyle Orland

–


|

19




Not so fast, GPT-6.1…


Credit:

Getty Images

Not so fast, GPT-6.1…


Credit:

Getty Images




Story text








OpenAI says it has canceled plans to release its updated GPT-6.1 model next month as it continues to investigate what testing shows to be a regression in terms of safety compared to previous models.

The move, first reported by The Wall Street Journal late Monday and later confirmed in OpenAI statements to the press, reflects what OpenAI Head of Safety Systems Saachi Jain said was a “trade off” between performance and security seen when testing the now-scrapped model. Jain said GPT-6.1 was better than previous models at sticking with difficult tasks all the way to completion without human intervention. But the model was also more likely to fail tests related to alignment (i.e. staying within the bounds set by human creators) and more willing to use sometimes “unsafe” tools and services to push ahead with a task. It was also more likely to try to deceive end users about actions it did or didn’t take, Jain said.

Last week, OpenAI said it was halting training of its “most capable models” following an incident where a model attempted to circumvent Internet access restrictions. GPT-6.1 was not among those “most capable models” covered by that move, OpenAI told the WSJ. And while GPT-6.1 won’t be released as is, the company said it intends to use the same base model for further training runs that it said will hopefully lead to future GPT-6 generation models.

The delay in GPT-6.1’s public release comes at a delicate time for OpenAI’s public safety reputation. Since the high-profile Hugging Face hacking incident this summer, OpenAI says it has notified dozens of third-parties about potential incidents caused by its models in testing. OpenAI said that includes stakeholders in “governments, universities, public agencies, and other institutions” and a breach of an Australian Medicare statistics site that drew a direct rebuke from the prime minister.

OpenAI was among a set of prominent AI companies publicly calling for a slowdown in model training and development over alignment concerns earlier this month. “When we talk about ‘pacing,’ we do not mean ‘stopping,’” OpenAI CEO Sam Altman said in a social media post this month. “Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs.”

While OpenAI was apparently uncomfortable with the security tradeoffs inherent to GPT-6.1 in its current state, similar tradeoffs are apparent in OpenAI’s current public models as well. A report released by the AI Security Institute on Monday found that GPT-6 was significantly more likely than previous GPT releases to perform “a range of unsanctioned attack activities” in simulated cybersecurity evaluations. Those “out-of-scope” actions include submitting malicious code to open source codebases and creating fake identities and benign code contributions to mask these actions.

Photo of Kyle Orland


Kyle Orland

Senior Gaming Editor
Kyle Orland has been the Senior Gaming Editor at Ars Technica since 2012, writing primarily about the business, tech, and culture behind video games. He has journalism and computer science degrees from University of Maryland. He once wrote a whole book about Minesweeper.


19 Comments

Leer artículo original en Ars Technica