VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

OpenAI GPT-6 Astra Saturates Key Benchmarks but Trails Fable 5.1 on Independent Coding Index

OpenAI launched GPT-6 Astra on September 3, 2026, scoring 99.9% on ARC-AGI-3 with a custom harness but trailing Anthropic's Fable 5.1 on independent…

OpenAI launched GPT-6 Astra on September 3, 2026, calling it a generational leap and declaring the AGI era has arrived8. The model is rolling out initially to a limited set of organizations and will become available over the coming days to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS2.

Astra is priced at $10 per million input tokens and $50 per million output tokens, matching the rate of Anthropic's Claude Fable 5 and 5.1. The API model label is gpt-6-astra.

OpenAI's self-reported benchmarks show Astra saturating several frontier evaluations. The model scored 97.6% on the hardest versions of FrontierMath and 99.9% on ARC-AGI-37. On security tasks, Astra scored 100% on ExploitBench (up from 78.5% for GPT-5.6 Sol), 42.4% on ExploitGym (Sol: 30.3%), and 99.2% within four attempts on SRE-Bench binary reverse engineering (Sol: 68.7%). On long-context retrieval, Astra achieved 100% on OpenAI's eight-needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens.

The ARC-AGI-3 result carries a significant caveat. The 99.9% score was achieved for $19K using OpenAI's custom Provider Adapter harness, which preserves opaque reasoning state between requests and uses compaction for longer conversations. Under the standard ARC-AGI harness, Astra scored 62.7% for $26K. For comparison, Claude Opus 5 scored 30.2% and GPT-5.6 Sol scored 7.8% on ARC-AGI-3.

Independent evaluations paint a more mixed picture. On the Artificial Analysis Intelligence Index, Astra scored 61, matching GPT-5.6 Sol's 61 and trailing Claude Fable 5.1's 661. On the Artificial Analysis Coding Agent Index, Astra scored 67, roughly equal to Claude Opus 5, Fable 5, and Muse Spark 1.3, but behind Fable 5.1's leading score of 703. Per task, Astra costs less than half the price of Claude Fable 5 for the same Coding Agent Index score.

ANALYSIS The gap between OpenAI's self-reported results and the independent Artificial Analysis scores is notable: Astra saturates OpenAI's chosen benchmarks yet ties or trails competitors on third-party indexes.

A persistent-agent variant, gpt-6-astra-aeon, has been confirmed as the name of a new long-running persistent agent4. Latent.Space, which received early access and reported burning over 20 billion tokens, described Astra as a fully capable AI engineer that can choose and train models, label data, keep pipelines saturated, instrument and read logs, deploy and debug entire systems, fan out and command subagents, and maintain coherence over billions of tokens in a single agent thread. The publication reported Astra ran at 33 tokens per second at the maximum $50 per million token rate, costing roughly $6 per hour. It spent about $100 over two days using Astra to monitor its own runs.

Latent.Space described Astra as the first Stargate and lightly looped supermodel from OpenAI. Greg Brockman said it is a generational leap and that OpenAI is now in the AGI era.