GPT-6 Astra: OpenAI's Bet That AI Should Do the Work
OpenAI's GPT-6 Astra, released September 3, 2026, shifts the frontier from answering questions to operating software. We break down the computer-use benchmarks, the safety controversy, and what the new pricing means for developers.
The largest model upgrade in OpenAI's recent history arrives with a clear thesis: the next generation of AI should not just talk about work — it should do it.
From Chatbot to Operator
GPT-6 Astra was released on September 3, 2026, and OpenAI describes it as its smartest, most human-intent-aligned model to date. It is also one of the company's largest training projects — the first model pre-trained at the Stargate Texas site on more than 100,000 GPUs, and the first to make heavy use of earlier models helping supervise the training run.
The headline capability is computer use. On the OSWorld 2.0 benchmark Astra scores 72.6%, ahead of GPT-5.6 Sol's 65.7%; on Agents' Last Exam it reaches 59.3% versus 55.5% for Claude Opus 5. In official demos it designed a PCB in KiCad, operated Excel and Power BI, built a house model in Blender and imported it into an interactive Unreal Engine 5 scene, and produced a complete piece of music from scratch through an MCP-connected Ableton.
The New Battlegrounds: Code and Science
OpenAI calls Astra its strongest software engineering model. On Terminal-Bench 4.0 it scores 57.7% against Sol's 37.3%, with lower estimated API cost per task; on DeepSWE v1.1 it reaches 74.1%. The Codex harness gained cross-context notes that survive beyond a single context window — it can now search earlier messages and tool outputs, and keep pushing forward on non-blocking tasks while waiting for the user to reply.
Science results are equally striking: 97.6% on FrontierMath Tier 4 and 96.0% on GPQA Diamond, and OpenAI reports Astra helped advance prime-gap research, narrowing a known upper bound on the distance between infinitely many pairs of primes.
Power, Guardrails, and the Price of Admission
Astra also crosses what OpenAI's Preparedness Framework defines as the critical cybersecurity threshold: 100% on ExploitBench versus 78.5% for Sol, and during evaluation it discovered and used two previously unknown zero-day vulnerabilities, which OpenAI disclosed to the maintainers. Access is tiered — vetted security teams enter through the Daybreak program, the general-release version refuses advanced cyber requests, and the API is priced at $10 per million input tokens and $50 per million output tokens.
President Greg Brockman suggested that years from now, Astra may be remembered as the start of the AGI era. The caveat: OpenAI acknowledges that Astra's reasoning is harder to monitor than Sol's, precisely because it finishes complex tasks with fewer wordy steps. The bottleneck is no longer raw capability — it is alignment.