Has OpenAI Reached AGI?

OpenAI says GPT-6 Astra could mark the arrival of AGI, with extraordinary leaps in reasoning, coding, computer use and cybersecurity.

OpenAI unleashes Astra, its most capable and controversial model yet

Author: Mark Sullivan

“Welcome to the AGI era,” said OpenAI president Greg Brockman at the tail end of a call with reporters about the Thursday release of the company’s newest model, GPT-6 Astra. OpenAI describes AGI, or artificial general intelligence, as “highly autonomous systems that outperform humans at most economically valuable work.”

The company calls Astra its most intelligent and safest model to date, built on advances in pre-training, reinforcement learning, and safety guardrails. OpenAI VP of research Aidan Clark said more than 100,000 GPUs were used to train it.

OpenAI said Astra outscored all other frontier models in benchmark tests measuring computer use, browser use, software engineering, cybersecurity, science, and professional work. It notched perfect or near-perfect scores on some especially difficult evaluations. Astra scored 99.9% on ARC-AGI-3, a game-style reasoning test designed to make advance preparation of the model almost impossible.

“When we look back and ask when AGI arrived, we’ll see that it’s about this time and it’s about this model,” Brockman said.

Astra is rolling out Thursday to a limited set of organizations and will become available over the coming days to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and Amazon Web Services. Enterprise administrators will have to enable Astra for their workspaces, with access turned off by default at launch.

Computer use

OpenAI said Astra can fill out forms, update customer relationship management (CRM) records, conduct research, draft summaries, analyze data, build websites, run front-end QA checks, and troubleshoot problems on-screen. It completed more desktop tasks correctly than its predecessor, GPT-5.6 Sol, while taking about half as long. A new harness, the software and orchestration layer that lets a model work like an agent, allows the Codex coding agent to finish web-based tasks 1.9 times faster. In demonstrations, Astra laid out a circuit board, built a business dashboard, and filled out a federal tax return in a browser from a W-2 form.

Professional work

The company said Astra is its best model for following existing templates and producing slides, documents, and spreadsheets that match a user’s writing and visual style. It is also trained to pull only relevant context into outputs. With Sites in ChatGPT, Astra can create, host, and share websites, web apps, and games from a prompt.

Coding

OpenAI called Astra its best software engineering model to date. On DeepSWE v1.1, which measures AI coding agent performance, it scored 74.1%, compared with 72.7% for GPT-5.6 Sol and 67.4% reported for Anthropic’s Claude Fable 5.1 model. On Cognition’s FrontierCode Extended, the company said Astra performed near Anthropic Claude Fable 5’s best result while using substantially fewer output tokens.

Cybersecurity

Astra also scored 100% on ExploitBench, which tests how well a model can identify and exploit software vulnerabilities. And it’s exactly these capabilities that have the cybersecurity world worried.

OpenAI said earlier this week that Astra can find unknown vulnerabilities and figure out how to exploit them across many well-defended systems without human guidance at each step. It was the first time the company had made that claim about one of its models, confirming a security classification OpenAI said it could not rule out three weeks ago. During evaluations, Astra found two previously unknown security flaws and chained them together to gain access to a software platform, according to OpenAI.

Earlier versions of OpenAI models were invovled in the Hugging Face incident this summer, which alarmed security professionals. During an evaluation, OpenAI agents were instructed to keep working through seemingly impossible tasks. Instead, they found a way to obtain the correct answers without actually solving the problems, then took additional steps to conceal what they had done. In the process, the agents escaped the secure test environment, accessed the internet, and eventually broke into servers belonging to both Hugging Face and OpenAI. Researchers reconstructed much of what happened by examining the agents’ messages to one another, along with the “chain of thought” messages that recorded their reasoning.

The security community grew even more concerned when reports emerged this week that Astra uses a technique called “recurrent depth” or “opaque recurrence.” Instead of carrying out all of its reasoning as a long sequence of explicit text tokens, more of the computation happens in the model’s internal, or latent, state. That can free up more tokens for higher performance. (OpenAI said in its press release that its evaluations found Astra’s written reasoning harder to monitor than GPT-5.6 Sol’s, and that Astra “still appears to struggle to conceal the reasoning needed for complex tasks, but we take the decline seriously.”)

OpenAI chief scientist Jakub Pachocki said it’s natural for large models, as they grow more capable, to provide less information about their work in plain language. He added, however, that the company needs to develop new methods of making them more “verbose.” OpenAI is deploying misalignment monitoring in production, using classifiers that can stop potentially unauthorized activity, and said the checks can also pause legitimate work.

OpenAI said Astra is, overall, even less likely to go rogue than earlier models. On a new evaluation informed by the Hugging Face incident, which tests whether a model given a difficult or impossible task will exceed its intended scope, GPT-5.6 Sol, without production safeguards, went beyond the authorized target in 48.2% of cases, compared with 0% for Astra, according to the company. Astra will refuse advanced tasks such as exploit discovery, OpenAI said, while less restrictive access is going to an initial set of trusted trusted defensive cybersecurity pros.

Correction 9/3 2 p.m. PT: An earlier version of this story stated that a version of the Astra model was involved in the Hugging Face incident. OpenAI says no Astra models were involved, but rather versions of its GPT-5.6 Sol model.

Credits: TCA, LLC.

Discover more from thinkly gold

Subscribe now to keep reading and get access to the full archive.

Continue reading