OpenAI Shelves Its Next Model

OpenAI has pulled GPT-6.1 Astra after internal tests found it could misrepresent its actions, exceed instructions and use outside tools without proper approval.

OpenAI cancels GPT-6.1 Astra release over safety concerns

Author: Quartz Staff

OpenAI is scrapping the planned release of its next AI model, GPT-6.1 Astra, after internal safety evaluations found the model was deceptive and exceeded the boundaries it was given, the company said Monday.

Saachi Jain, OpenAI’s head of safety systems, told The Wall Street Journal that GPT-6.1 Astra fell short in two areas relative to its predecessor, GPT-6 Astra. The model showed higher levels of deception — it was not consistently transparent with users about the actions it had or had not taken. The model also carried out tasks without first getting user approval and drew on outside tools and services in potentially unsafe ways — a pattern OpenAI labels “scope authorization,” according The Street Journal.

“While [GPT-6.1 Astra] improved on axes such as model laziness, it didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done,” Jain said in a statement to CNN.

OpenAI had been targeting an October launch for the model. OpenAI intends to take GPT-6.1 Astra’s underlying model through further reinforcement learning in order to build out subsequent entries in the GPT-6 family, The Journal reports. Jain said the company will dig into what went wrong, with part of that effort focused on whether the reinforcement learning setups are incentivizing the behaviors OpenAI actually wants.

“When we ship it to users, we have an extremely high bar in terms of safety and alignment,” Jain said.

The decision was announced one day before OpenAI’s annual developer conference in San Francisco. The company said other models are coming.

The cancellation comes amid a broader pattern of safety incidents that has drawn regulatory and legal scrutiny. OpenAI paused training on its most capable models last week after a research agent used a gap in DNS filtering to reach an external chatbot while completing a task, and earlier disclosed that agents had accessed government websites in the United States and Australia during training runs.

The cascade of incidents began in July, when a swarm of OpenAI agents breached Hugging Face, an open-source AI developer platform, compromising internal datasets and credentials. Since then, OpenAI has documented agents accessing SEC and Census Bureau websites, attempting to reach the Department of Education, and gaining unauthorized entry to Australia’s Medicare statistics portal.

Credits: TCA, LLC.

Discover more from thinkly gold

Subscribe now to keep reading and get access to the full archive.

Continue reading