OpenAI released GPT-6 Astra Thursday, calling it its most capable and most aligned model yet, and within hours the launch had turned into two separate stories, one about genuinely record setting benchmark results, and another about a rollout messy enough that CEO Sam Altman was publicly apologizing to paying customers by the next morning.
What Astra actually does differently
OpenAI’s own results show Astra topping a wide range of technical benchmarks, scoring 98% on FrontierMath Tier 4 and reportedly helping establish new mathematical results on the gaps between prime numbers, including improving a bound that had stood for more than 80 years. On professional and computer use tasks, Astra scored 59.3% on Agents Last Exam and completed OSWorld 2.0 tasks in roughly 40 minutes on average, well ahead of its predecessor GPT-5.6 Sol. The model also posted a perfect 100% on ExploitBench, a cybersecurity benchmark measuring whether a model can turn known vulnerabilities into working exploits, a capability OpenAI itself says now meets the critical threshold under its internal risk framework.
The gap between the headline number and the independent verification
OpenAI’s launch page states plainly that Astra saturates ARC-AGI-3 with a 99.9% score. The organization that actually runs that benchmark, the ARC Prize Foundation, tells a more layered version of the same result. Under its Standard harness, the same minimal interface used to compare every model on equal footing, Astra scored 62.7%, a strong result but nowhere near saturation. The 99.9% figure only appears under a separate Provider Adapter harness that lets Astra preserve hidden reasoning state and compress long conversations using tools specific to its own provider, conditions ARC Prize is careful to label separately from the apples to apples comparison. Both numbers are real, but only one of them describes how Astra performs under the same rules every other model was tested against.
A rollout messier than the benchmarks suggest
The technical results arrived alongside a launch that frustrated OpenAI’s own paying subscribers. Astra rolled out first to a limited set of enterprise customers with access to OpenAI’s Daybreak cybersecurity platform, with general availability for Plus, Pro, Business and Enterprise users promised only in the days that followed. Subscribers on the higher priced Pro tier, who are accustomed to first access at launch, reacted with visible frustration on social media, prompting Altman to post an apology acknowledging the rollout had been messy and asking for patience without offering a specific timeline for full access.
The safety stakes behind the numbers
Astra’s cybersecurity capabilities are a central part of both the excitement and the concern surrounding this release. Beyond the perfect ExploitBench score, OpenAI reported that Astra discovered two previously unknown zero-day vulnerabilities during testing against recent software flaws, which the company says it is disclosing to the affected maintainers. OpenAI frames this as evidence Astra can help defenders find and patch weaknesses faster, and says the version launching today will refuse more advanced requests like building proof of concept exploits, with less restrictive access planned later through Daybreak. At the same time, independent reporting has referenced researcher concern about safety risks tied to this release, a thread not fully detailed in the materials reviewed for this piece but consistent with the scrutiny any model crossing this kind of capability threshold tends to draw.
What Astra is, and what OpenAI says it isn’t
OpenAI is careful to frame Astra’s results as meaningful progress rather than a finished destination, and outside evaluators have echoed that framing. The ARC Prize Foundation, whose own benchmark exists specifically to measure the distance between current AI and general intelligence, described Astra as a noticeable step change in frontier capability while explicitly declining to call saturating its test proof of achieving AGI. That caveat sits somewhat awkwardly next to OpenAI’s own marketing language describing Astra as the start of the AGI era, a gap between technical restraint and promotional framing that ran through nearly every part of this launch, from the benchmark table down to the release schedule itself.

