GPT-6 Astra vs. Fable 5.1 benchmarks: which scores are comparable and which aren't
On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.
We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.
OpenAI announced GPT-6 Astra on September 3, 2026, but it was not yet publicly available as of September 4. Initial access was limited to selected companies in OpenAI’s Daybreak cybersecurity program, with paid ChatGPT plans, the API, and AWS due to follow “in the coming days.”
Astra’s strongest reported gains are in science, visual engineering, computer use, and long-running agent tasks. The coding results do not establish a clear leader, and several comparisons with Anthropic’s Fable 5.1 use different test versions or settings. This guide separates those reported numbers from what they can actually prove.
Latest update, September 4, 2026: OpenAI has published API documentation for standard GPT-6 Astra, including its context window, prices, endpoints, and model ID. Public access is still described as coming in the following days. No separate Astra Pro developer page or API model ID was found in the official documentation checked.
GPT-6 Astra at a glance
| Question | Answer |
|---|---|
| What is GPT-6 Astra? | OpenAI’s new flagship GPT-6 model for complex agent, research, coding, and computer-use tasks. |
| Is there an Astra Pro model? | Yes. OpenAI announced GPT-6 Astra and GPT-6 Astra Pro. Pro access is planned for Pro, Business, and Enterprise users. |
| What is Astra’s context window? | OpenAI documents a 1,050,000-token context window, with up to 922,000 input tokens and 128,000 output tokens. |
| How much will the Astra API cost? | OpenAI announced pricing of $10 per million input tokens and $50 per million output tokens once Astra becomes available through the API. |
| Is Astra better than Fable 5.1? | It leads some vendor-reported evaluations, but there is no clean overall winner. The models were not always tested with the same benchmark version, harness, reasoning setting, or access level. |
GPT-6 Astra release date and public availability
The date commonly reported as the GPT-6 Astra release date is September 3, 2026, when OpenAI announced the model and began a restricted rollout. The first users were a limited group of Daybreak companies, not general ChatGPT or API customers.
OpenAI said Plus, Pro, Business, and Enterprise users, the API, and AWS would get Astra “in the coming days”; TechCrunch described the window as the following week. No account-by-account schedule was published. The New Stack reported that eligible API customers would receive Zero Data Retention support when access begins.
That wording matters if you are searching for Astra in ChatGPT or the API and cannot see it. Headlines use “launched” and “released,” but those terms describe the announcement and restricted first access, not public availability. There was no confirmed free-tier rollout, and OpenAI had not said whether free ChatGPT users would receive Astra.
The announced lineup has two variants:
- GPT-6 Astra, the flagship model.
- GPT-6 Astra Pro, an expanded option planned for Pro, Business, and Enterprise accounts.
There were no announced GPT-6 equivalents of the GPT-5.6 Luna, Terra, and Sol variants as of September 4.
Is GPT-6 Astra Pro officially documented?
OpenAI announced GPT-6 Astra Pro for Pro, Business, and Enterprise users, which explains the growing searches for an Astra Pro model page and API identifier. The documentation picture was incomplete when checked on September 4:
- Standard Astra has an official developer page and the API model ID
gpt-6-astra. - OpenAI had not published a separate Astra Pro API model ID, price, context window, or endpoint table in the sources checked.
- The OpenAI model release notes did not yet contain an Astra Pro entry.
This does not mean Astra Pro was cancelled or that it will never receive API access. It means the launch announcement named the model and eligible ChatGPT plans before complete public developer documentation appeared. Treat model IDs inferred from naming conventions as unverified until OpenAI lists them.
GPT-6 Astra benchmark results
These are the main Astra benchmark scores reported at launch. They should be read as a map of where the model appears strong, not as one combined ranking.
Direct same-run gain Promising, with caveats Context only
| Benchmark | Astra | Launch comparison | Readout |
|---|---|---|---|
| DeepSWE v1.1 | 74.1% | Sol 70.8%; Fable 5.1 67.4%; Muse Spark 1.3 75.4% at maximum reasoning | Clear gain vs SolTop coding results remain close. |
| ARC-AGI-3 | 98.6% | No direct Fable result cited | System scoreIncludes OpenAI's harness. |
| FrontierMath Tier 4 | 97.6% | No direct Fable result cited | Private tasksIndependent replication is limited. |
| BenchCAD Vision2Code | 95.9% | Fable 5.1 84.3%; Sol 83.3% | Clear gain vs SolFable used modified settings. |
| Terminal-Bench Science | 64.6% | Fable 5.1 52.6% | Reported leadExact run parity is unconfirmed. |
| OSWorld V2-Offline | 72.6% | Sol 65.7% | Clear gain vs SolAverage task time also fell. |
| ExploitGym | 42.4% | Sol 30.3% | Qualified gainThe usual six-hour limit was removed. |
| ExploitBench | 100% | No broad comparison available | Restricted contextAccess controls affect real use. |
Astra’s coding benchmark is good, not dominant
On DeepSWE v1.1, Astra scored 74.1% across 113 agentic coding tasks, up from the 70.8% OpenAI reported for GPT-5.6 Sol in its launch evaluation. Meta reported 75.4% for Muse Spark 1.3 at maximum reasoning, while a separate public leaderboard run put Gemini 3.8 Flash and Claude Opus 5 around 74% and Sol around 73%. Those Sol figures come from different runs and should not be treated as contradictory. Muse 1.3 was available at launch, but its maximum reasoning setting remained under safety review and was not generally available. The uncertainty ranges overlap, and 1.3 percentage points on this test represents roughly one or two tasks.
OpenAI’s own chart put Fable 5.1 at 67.4%. That gives Astra a visible lead in that specific run, but it should not be generalized to every repository, coding harness, or software-engineering task.
The defensible conclusion is that Astra improves on OpenAI’s previous flagship and appears competitive with the top group. Teams should still test it on their own issue set before making a purchasing decision. Our coding-agent validation guide explains how to build a representative task set.
The biggest gains are in science and engineering
The model scored 95.9% on the 1,000-file Vision2Code subset of BenchCAD. That benchmark asks a model to recreate CAD programs from rendered views and measures the geometric overlap of the resulting 3D models. OpenAI reported 84.3% for Fable 5.1 and 83.3% for Sol. The comparison with Fable used modified evaluation settings, but the gain over OpenAI’s own previous model is substantial.
On Terminal-Bench Science, Astra reached 64.6%, compared with Anthropic’s reported 52.6% for Fable 5.1. This test covers 70 command-line research tasks across five scientific fields. Here again, a vendor result and a public leaderboard result are not automatically equivalent unless the harness, task set, time limit, and scoring procedure match.
The 97.6% FrontierMath Tier 4 score deserves both attention and context. Epoch AI says Tier 4 contains 43 difficult private problems, while the reported Astra result appears to cover the 41 problems available to OpenAI. Epoch also discloses that OpenAI funded the benchmark and has exclusive access to part of it. That does not invalidate the result, but independent replication is limited by design.
ARC-AGI-3 tests the system, not only the model
Astra’s 98.6% ARC-AGI-3 result is likely to attract the largest headlines. It is not a model-only score.
OpenAI ran Astra through a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts. OpenAI has separately shown that changing these system settings can sharply increase ARC-AGI-3 performance without changing the underlying model.
The score therefore tells us something useful about the Astra agent system. It does not isolate how much of the gain comes from model weights, reasoning effort, retained state, compaction, or the surrounding harness. That distinction matters when comparing Astra with a model running inside another product.
GPT-6 Astra vs Fable 5.1
There is no single reliable leaderboard for “Astra vs Fable 5.1.”
Winner in this run Reported lead, settings differ Tie Not comparable
| Category | GPT-6 Astra | Fable 5.1 | Winner |
|---|---|---|---|
| Agentic coding | 74.1% DeepSWE v1.1 |
67.4% OpenAI's chart |
AstraWinner in OpenAI's run; harnesses still matter. |
| CAD reconstruction | 95.9% Vision2Code |
84.3% OpenAI's comparison |
Astra reported leadFable used modified settings. |
| Scientific terminal tasks | 64.6% | 52.6% | Astra reported leadExact run parity is not established. |
| Desktop computer use | 72.6% OSWorld V2-Offline |
77.9% Different OSWorld release |
No winnerBenchmark versions differ. |
| API price | $10 / $50 per 1M input / output tokens |
$10 / $50 per 1M input / output tokens |
TieCost per completed task is unknown. |
| Context window | 1.05M 922K max input; 128K max output |
Not disclosed on the cited product page | No comparisonOnly Astra has a confirmed figure here. |
| Persistent agent work | Searchable notes; asynchronous clarification | Different product and harness features | Test directlyFeature names do not show reliability. |
The clearest Astra advantages in the launch data are BenchCAD and Terminal-Bench Science. Fable 5.1 has a higher published OSWorld figure, 77.9% versus 72.6%, but Anthropic says its result uses a different release and should not be compared with earlier scores.
Price also fails to break the tie. Both companies list $10 per million input tokens and $50 per million output tokens for these models. A real cost comparison needs total tokens, retries, tool calls, latency, and successful completion rate on the same tasks.
What developers are sharing: games, limits, and computer use
Early discussion around Astra and Fable 5.1 is being driven as much by demos and subscription limits as by benchmarks. The posts reveal what developers want to test, but not how reliably the models perform.
Fable 5.1 one-shot games
An X trend about Fable 5.1 collected early demonstrations of playable racing games generated from short prompts and a first-person game reconstructed from a screenshot. The page also cited a much longer, roughly 30-hour reconstruction run, which is materially different from a literal one-prompt, one-response result.
These are compelling zero-to-one prototypes, not evidence of production readiness. Frame rate, collision edge cases, mobile controls, asset licensing, accessibility, maintainability, and stability after revisions still need testing. The Fable 5.1 launch discussion provides more examples; a reproducible one-prompt claim should include the prompt, files, settings, elapsed time, interventions, and revision history.
The Fable 5.1 usage-limit debate
Some Claude subscribers report exhausting their rolling allowance unusually quickly during long Fable 5.1 coding sessions, including on high-priced plans. The X trend summary and this video discussion of Fable 5.1 limits reflect that concern. Claims that a five-hour allowance can disappear in under 30 minutes remain anecdotal: usage varies with plan, context, tool calls, reasoning effort, service capacity, and repository size.
There is also a distinction between subscription limits and API economics. Anthropic’s lower cache-read price can reduce metered API costs for repeated context, but it does not give Claude subscribers a 75% larger message allowance. For agentic work, measure completed tasks per allowance or per dollar rather than messages alone.
Astra computer-use demonstrations
OpenAI’s launch examples focused on delegated desktop work. Reuters reported that OpenAI showed Astra completing cat-sitter research in 5 minutes 27 seconds, versus a stated 30-minute human baseline, and a job search in 2 minutes 51 seconds, versus a stated five-hour baseline. Reuters also listed tax preparation, game development, architectural rendering, legal memo formatting, and apartment hunting among OpenAI’s examples.
These were company demonstrations, not independent benchmarks. Results depend on task scope, completion criteria, available websites, excluded login or CAPTCHA steps, and how the human baseline was measured. General users could not yet reproduce them.
Visual demos showing Astra generating SVG interfaces or interactive 3D scenes are best read the same way. They can reveal breadth and speed, but visual polish does not establish semantic HTML, responsive behavior, accessibility, browser compatibility, or maintainable code. The linked computer-use walkthrough and visual generation demo are starting points for inspection, not independent proof of “superhuman” computer use.
All three trends are about time to a convincing first result: Fable 5.1 is attracting attention for ambitious prototypes, while Astra is being presented as unusually fast at operating existing software. Reliability, hidden intervention, and total cost remain open questions.
What changes for Codex and coding agents?
Astra’s most practical developer feature may be how it handles work that exceeds one context window. Compaction can discard details an agent later needs, such as a failed approach, test command, file path, or early constraint. OpenAI says Astra can instead keep searchable notes across context windows. The feature is experimental behind a config.toml setting and is expected to become the Astra default in Codex in the coming weeks. That surrounding system matters, as our guide to coding-agent harnesses explains.
Astra can also ask a clarifying question while continuing work that does not depend on the answer. That sounds small, but it changes the control flow of an agent run: one blocked decision no longer has to stall every independent task.
OpenAI also demonstrated Astra using KiCad, Excel, Blender, Power BI, browsers, and website QA workflows. On Mind2Web, OpenAI says Astra with the updated Codex harness completed tasks 1.9 times faster than the existing Sol-based configuration. Because both the model and harness changed, the result cannot attribute the speedup to Astra alone.
GPT-6 Astra pricing and API cost
OpenAI announced this pricing for future Astra API access:
- Input: $10 per million tokens.
- Cached input: $1 per million tokens.
- Cache writes: $12.50 per million tokens.
- Output: $50 per million tokens.
That matches Fable 5.1’s listed price and is 2.5 times GPT-5.6 Sol’s promotional rate cited in launch reporting. General API customers could not yet verify Astra billing in production.
OpenAI’s model page adds two cost details that matter for long agent runs. Requests with more than 272,000 input tokens are charged at twice the input and cache rates and 1.5 times the output rate for the full request. Batch and Flex processing are listed at half the standard rates, while Fast mode costs twice the applicable rate.
Per-token price is not the same as cost per task. A more expensive model can be cheaper if it finishes in fewer steps, uses fewer retries, or avoids human cleanup. OpenAI says Astra uses fewer tokens in several evaluations and partner tests, but the launch evidence is not detailed enough to calculate a reliable cost-per-completed-task advantage. See how to track token ROI for a fuller cost framework.
For an internal evaluation, track:
- Successful tasks, not attempted tasks.
- Input, output, cached, and tool-related token costs.
- Wall-clock latency and rate-limit delays.
- Number of retries or human interventions.
- Review time and defect rate after completion.
Those measurements will tell you more than multiplying one sample prompt by the list price.
Safety, cybersecurity, and access limits
OpenAI says Astra crossed the “Critical” cybersecurity threshold in its Preparedness Framework. In company testing, it developed exploits for hardened browsers and operating systems and found two previously unknown vulnerabilities while being evaluated against recent V8 bugs. ExploitBench measures progress from reaching vulnerable code through exploit primitives to arbitrary code execution. OpenAI said it was disclosing the newly found issues to maintainers.
The reported cyber results reflect Daybreak Blue access, not the production configuration planned for broader access. Daybreak Blue is a vetted-access program for authorized defensive work; it is not a separate Astra model. OpenAI says standard access will refuse some advanced cybersecurity tasks, including exploit discovery.
OpenAI also reported a harder monitoring problem. Astra’s written reasoning was less observable than Sol’s in tests intended to elicit monitoring evasion. That disclosure is a useful counterweight to the company’s description of Astra as its most aligned model. Greater capability and easier oversight are not the same trend.
Future API users should also expect cybersecurity safety checks that may stop tasks rather than pause for approval. In its launch disclosures, OpenAI warned that these checks could sometimes slow or block unrelated work.
How to evaluate Astra for your own work
Do not reproduce a public benchmark badly. Build a small evaluation around the work you actually need the model to do.
- Select 20 to 50 representative tasks with known acceptance criteria.
- Run Astra and Fable 5.1 through comparable harnesses, tools, permissions, and time budgets.
- Keep reasoning effort and retry policy fixed where the APIs allow it.
- Score task completion blind before comparing style or preference.
- Record cost, latency, intervention, and review cleanup alongside accuracy.
- Repeat enough runs to expose variance instead of trusting one successful demo.
For coding agents, include repository navigation, a bug fix, a multi-file change, a test failure, and a task that should trigger clarification. Persistent memory is only valuable if it preserves the right constraint and retrieves it at the right time.
The bottom line
GPT-6 Astra posts strong vendor-reported results in CAD reconstruction, scientific terminal work, advanced math, computer use, and long-running agents. It improves on GPT-5.6 Sol, but the data does not establish a universal winner over Fable 5.1: coding results remain close to the broader field, OSWorld versions differ, and some headline scores use private tasks or measure a model-plus-harness system.
The reason to test Astra is its combination of competitive capability and searchable state across context windows. Real repository performance matters more than launch charts.
Frequently asked questions
When did GPT-6 Astra come out?
OpenAI announced GPT-6 Astra and began restricted Daybreak access on September 3, 2026. This is the date commonly described as its launch or release.
Is GPT-6 Astra available in ChatGPT?
Not yet as of September 4, 2026. OpenAI said Astra would roll out to Plus, Pro, Business, and Enterprise users “in the coming days,” but it did not publish an exact schedule. The company did not announce availability for free ChatGPT accounts.
Is GPT-6 Astra better than Claude Fable 5.1?
Not conclusively. Astra leads several vendor-reported science, CAD, and coding runs, but different benchmark versions and harnesses prevent a universal ranking; test both models on your own tasks.
What did GPT-6 Astra score on DeepSWE v1.1?
OpenAI reported 74.1% across 113 tasks. That beat Sol’s 70.8% in OpenAI’s launch evaluation and Fable 5.1’s 67.4% in the same chart. Meta separately reported 75.4% for Muse Spark 1.3 at a maximum reasoning setting that was not generally available.
What is GPT-6 Astra’s ARC-AGI-3 score?
OpenAI reported 98.6%, but the result includes a Responses API harness that retains reasoning and uses compaction. It is best understood as an Astra-system score rather than a clean model-only comparison with Fable 5.1.
Does Claude Fable 5.1 have an ARC-AGI-3 score?
Anthropic did not publish an official Fable 5.1 ARC-AGI-3 score in the launch materials checked. ARC-AGI-3 is not exclusive to OpenAI, but Astra’s reported 98.6% result used an OpenAI Responses API harness, so it cannot stand in for a matched Astra-versus-Fable comparison.
What did Astra score on Terminal-Bench Science 0.1?
OpenAI reported 64.6%, compared with Anthropic’s 52.6% for Fable 5.1. Exact run parity has not been established, so the numbers indicate a reported Astra lead rather than a fully controlled head-to-head result.
What is GPT-6 Astra’s BenchCAD score?
Astra scored 95.9% on BenchCAD’s 1,000-file Vision2Code subset with Python tools. OpenAI reported 84.3% for Fable 5.1 using modified evaluation settings and 83.3% for GPT-5.6 Sol.
Can Claude Fable 5.1 build a game from one prompt?
Early social demos show Fable 5.1 producing playable game prototypes from short prompts. “One prompt” does not necessarily mean a production-ready game or zero human intervention: check whether the creator published the prompt, files, settings, elapsed time, and revision history.
Why does Fable 5.1 hit usage limits so quickly?
Long coding runs can consume substantial context and reasoning tokens, and some subscribers report reaching rolling usage limits quickly. There is no universal 30-minute limit: usage varies by plan, repository size, reasoning effort, tool activity, message length, and service capacity.
Is Astra better than a human at computer use?
OpenAI demonstrated Astra completing selected research and job-search tasks much faster than stated human baselines, but these were company examples rather than independent tests.
What is the GPT-6 Astra API price?
OpenAI announced a future API price of $10 per million input tokens and $50 per million output tokens, matching Fable 5.1’s listed per-token price.
What is Astra Pro?
GPT-6 Astra Pro is the higher-tier model announced alongside Astra for Pro, Business, and Enterprise users. As of September 4, OpenAI had not published a separate developer model page, API model ID, price, or context-window specification for Astra Pro.
Is GPT-6 Astra Pro available through the API?
No public Astra Pro API model was documented when checked on September 4. OpenAI’s standard Astra page lists gpt-6-astra, but the official documentation checked did not list a corresponding Astra Pro model ID. Do not send production requests using an inferred ID until OpenAI documents one.
What is the GPT-6 Astra context window?
OpenAI documents a 1,050,000-token context window for standard GPT-6 Astra, including a maximum of 922,000 input tokens and 128,000 output tokens. Requests above 272,000 input tokens use higher long-context pricing.
Does Astra have persistent memory?
OpenAI says Astra can keep notes across context windows and search earlier messages and tool output in Codex. The feature was experimental at launch and is expected to become the Astra default in the coming weeks. It is persistent task state inside the agent workflow, not a claim that every Astra product remembers every user conversation indefinitely.
Source: DevAgentStack · Field Notes · devagentstack.com