In the first week of September 2026 the two biggest AI labs shipped their biggest models, three days apart. Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1. OpenAI released GPT-6 Astra to approved users on September 3 and to everyone on paid plans the next day, with its president calling it the start of the AGI era. The two models cost the same to developers, split the independent scoreboards, and win different jobs. We read the primary sources, ran both models on the same four everyday tasks, and wrote the comparison for anyone deciding which one to open tomorrow morning, not just for engineers.
Key Takeaways
- Claude Fable 5.1 is the better everyday assistant for writing, long documents, and careful analysis; GPT-6 Astra is the better pick for math, for agents that operate your computer and browser, and for short high-volume tasks. Neither is a clean sweep.
- The scoreboards disagree. Artificial Analysis ties them at 53 on its Intelligence Index, LMArena's agent board puts Fable 5.1 first inside overlapping error bars, and Epoch's index and the hardest math tests favour Astra. Every "Astra beats Fable" table you have seen came from a third party, because neither lab benchmarks the other.
- For developers both list at $10 per million input tokens and $50 per million output tokens, but Fable 5.1 reads cached context at a quarter of Astra's price and Astra doubles its rate above 272,000 tokens; for individuals, neither model is on a free tier, ChatGPT Plus is $20 a month, and Claude Pro starts at $17 a month on an annual plan but reaches Fable only through usage credits.
- Astra is the first OpenAI model rated Critical for cybersecurity under OpenAI's own rules, ships with reasoning you cannot read, and went through a voluntary US government review; Anthropic keeps the unlocked version of its model, Mythos 5.1, behind an invitation-only program for vetted US organizations.
- In our identical four-task test, neither model invented a source, both caught all three errors we planted in a spreadsheet, Fable 5.1 wrote the better opening, and Astra mistook a year for an ID. Fable 5.1 wrote 35% more text for the same questions.
- Our predictions, with dates: Mythos stays locked through 2026, OpenAI ships a cheaper or tiered Astra by year end, and the list price does not move.
Which is better, GPT-6 Astra or Claude Fable 5.1?
Direct answer: As of September 9, 2026, Claude Fable 5.1 is the better default for writing, reading long documents, and analysis you have to defend, and GPT-6 Astra is the better choice for math, for agents that fill forms and work inside apps, and for short tasks at volume. Independent indices tie them overall. Both charge developers $10 and $50 per million tokens; Fable 5.1 is cheaper on cached context, Astra is cheaper per completed task because it writes less.
What just happened, in plain English?
Anthropic's model comes in two versions with two names. Fable 5.1 is the one anyone can pay for. Mythos 5.1 is, in Anthropic's own words, "the same model, but with different levels of safeguards," and it is "available only through our trusted access programs" for vetted cybersecurity and life-sciences organizations in the US (Anthropic). OpenAI's model has one name, but its system card discloses that Astra "meets our Critical threshold" for cybersecurity, the highest level in OpenAI's Preparedness Framework, and that the first customers were its Daybreak cybersecurity program before the general rollout (GPT-6 Astra System Card). OpenAI president Greg Brockman told Fortune "it's not unreasonable to feel that we are now in the AGI era" (Fortune).
| Date | What happened | Source |
|---|---|---|
| Sep 1 | Claude Fable 5.1 (for everyone on paid plans and the API) and Claude Mythos 5.1 (invitation only) released; cache reads cut 75% to $0.25 per million tokens | Anthropic model page |
| Sep 3 | GPT-6 Astra to approved organizations, starting with Daybreak cybersecurity customers; system card published; general availability in Microsoft Foundry | TechCrunch, Microsoft |
| Sep 4 | Astra reaches ChatGPT Plus, Pro, Business and Enterprise, the API, and AWS | Wikipedia, GPT-6 Astra |
| Sep 8 | Astra generally available on Amazon Bedrock | AWS |
| Sep 29 | OpenAI DevDay 2026 in San Francisco, the next scheduled OpenAI stage | devday.openai.com |
Which one should you use for what?
This is the table most people came for. It combines the independent benchmarks, both vendors' own tables, our four-task test, and the price mechanics explained further down. Where a row rests on one run of one prompt, it says so.
| If you mostly... | Pick | Why |
|---|---|---|
| Write: emails, essays, posts, anything a person will read | Fable 5.1 | Practitioners in Zvi Mowshowitz's roundup call it "more incisive, improved writing" with fewer Claude mannerisms; in our test its opening was the one an editor would keep |
| Code | Fable 5.1 for long sessions, Astra for terminal-heavy automation | Fable 5.1 leads the Artificial Analysis Coding Agent Index 70 to 67 as quoted by DataCamp; Astra leads Terminal-Bench on every harness that scored both |
| Study, research, and check facts | Fable 5.1, narrowly | Fable 5.1 takes Humanity's Last Exam with tools, 65.0% to 57.2%; Astra takes GPQA Diamond, 96.0% to 93.7% (DataCamp). Neither invented a source in our test |
| Do math and statistics | Astra | 97.6% on FrontierMath Tier 4 against 87.8%, and 2 of 68 FrontierMath Erdos problems solved against 0 (Zvi) |
| Work in spreadsheets and exports | Either; Fable for the write-up, Astra for the quick check | Both caught all three errors we planted; Fable 5.1 asked more questions, at 2.2x the cost, in one run |
| Let an agent fill forms, click through sites, and use apps for you | Astra, with a caveat | OpenAI's tagline is "Anything you can do on a computer, Astra can do for you. Fast." and it scores 92.7% on ScreenSpot-Pro (OpenAI, MindStudio). Claude in Chrome went fully autonomous on August 26 on every paid plan (Anthropic) |
| Feed it very long documents, whole archives | Fable 5.1 | Both hold about a million tokens, but Astra doubles its input price above 272,000 tokens and Fable does not |
| Want answers fast | Fable 5.1 | 68 tokens per second against 54 on Artificial Analysis's measurement |
| Need to explain to someone why the AI said what it said | Fable 5.1 | Fable returns a readable summary of its reasoning; Astra's comes back encrypted, and OpenAI says the model is harder to monitor than its predecessor |
| Run marketing or data work with real company data | See the work section below | The model matters less than whether it can see your data |
What does each one cost a normal person?
Neither flagship is free. OpenAI's announcement lists Pro, Enterprise and Business Premium plans for immediate access, with Plus and Business "rolling out over the coming days" (OpenAI); ChatGPT Plus is $20 a month, which Zvi's roundup notes is the cheapest door to Astra, against the $100 Claude Max plan he cites as the comfortable way to run Fable 5.1 (Zvi). On the Claude side, Pro costs $17 a month on an annual plan or $20 monthly and reaches Fable only through usage credits; Max starts at $100 a month and caps Fable at "50% of weekly limits" (claude.com). In plain terms: on a $20 plan you will get Astra with message caps, and you will get Fable 5.1 only while your credits last. If you use either model for hours a day, the API is cheaper than the top consumer tiers.
For developers, the two carry the same sticker and a different bill. Both publish $10 per million input tokens and $50 per million output tokens (OpenAI, Anthropic). Four things differ. Fable 5.1 reads cached context at $0.25 per million tokens, 2.5% of its input price, while Astra charges $1.00, and other Claude models pay 10% (Anthropic models overview). Astra charges $20 for input and $75 for output once a prompt passes 272,000 tokens; Fable 5.1 documents no long-context surcharge (OpenAI pricing). Astra offers batch processing at half price and a fast mode at double; Fable 5.1's batch rate is listed at the same $5 and $25 on OpenRouter. And Fable 5.1 writes more: Anthropic's own claim is that it costs "25% less than Fable 5 for typical workloads," but R&D World recalculated that it uses about 1.7 times the output tokens of Fable 5 and costs about 20% more per task, and Artificial Analysis's cost-per-task metric lands at $3.26 for Astra and $7.63 for Fable 5.1 at maximum effort (Artificial Analysis leaderboard).
Two worked examples. A team running 300 report requests a month, each with 40,000 tokens of cached context, 2,000 fresh input tokens and 1,500 output tokens, pays about $55.50 on Astra and $46.50 on Fable 5.1 at equal output, counting fresh input, daily cache writes, cached reads and output together; the whole difference is the 12 million cached tokens, which cost $12 on one and $3 on the other. Feed one 400,000-token archive into a single request with a 5,000-token answer and Astra bills roughly $8.40 to Fable's $4.25. Then add the verbosity: with the 35% extra output we measured, the first example lands within a dollar; with R&D World's 1.7x, Astra is about 12% cheaper. A usable rule of thumb from the rate cards: Fable 5.1 is cheaper once you read about 20 cached tokens for every token it writes, which is what long-context, short-answer work looks like; below that, Astra is. The bigger lever on both sides is the effort dial: on Artificial Analysis both keep their score of 53 one effort notch below maximum, while cost per task drops 29% for Astra and 22% for Fable 5.1.
Who actually wins the benchmarks?
Depends on whose board you read, and the honest answer is that the gap between these two is smaller than the gap between either and the task you actually have. Four things the primary sources show.
Neither lab benchmarks the other. OpenAI's announcement claims state-of-the-art results on Agents' Last Exam, AutomationBench, ScreenSpot Pro, FrontierMath Tier 4, ARC-AGI 3, TerminalBench-4.0, Terminal-Bench Science and HealthBench Pro, and prints no Claude number; its system card compares only to GPT-5.6 Sol (OpenAI, system card). Anthropic's table compares Fable 5.1 to Fable 5, Opus 5 and GPT-5.6 Sol, a model two months older than Astra.
The independent boards split. On the live Artificial Analysis comparison on September 9, Astra and Fable 5.1 tie at 53 on the Intelligence Index; Fable 5.1 leads AA-Briefcase (1,662 to 1,562), GDPval-AA v2 (1,764 to 1,580) and SciCode (63% to 56%), while Astra leads AutomationBench-AA (68% to 59%) and Terminal-Bench v4.0 (59% to 52%). Launch-week snapshots quoted by DataCamp (66 to 61) and Zvi (57 to 55 on the revised index) had Fable 5.1 ahead, and Zvi reports Epoch's Capabilities Index at 169 for Astra against 163 for Fable 5.1. LMArena's agent leaderboard, the one board that measured both under one harness the day we checked, ranks Claude Fable 5.1 (Max) first at 14.51% and GPT-6 Astra (Max) second at 12.55%, with error bars that overlap (arena.ai).
Astra owns the math and the screen; Fable owns knowledge work. Third-party tables agree on direction: Astra 97.6% to 87.8% on FrontierMath Tier 4 and 96.0% to 93.7% on GPQA Diamond (DataCamp); Fable 5.1 65.0% to 57.2% on Humanity's Last Exam with tools, and 1,853 on GDPval-AA v2 on Anthropic's own table against 1,824 for Opus 5. Anthropic also reports Fable 5.1 more than doubling Fable 5 on Terminal-Bench-Science, 52.6% against 24.7%, with a stated error of 3.5 to 4.5 points (Anthropic).
The headline numbers depend on the test setup. Astra's widely quoted 99.9% on ARC-AGI-3 came from OpenAI's own configuration; the standard harness used for rival models put it at 62.7% (The New Stack), and Fortune reports 66% without tools against 7.8% for GPT-5.6 Sol (Fortune). On Anthropic's side, Mythos 5.1 scores 60.9% on Terminal-Bench 4.0 and Fable 5.1 55.8% on the same table despite being the same model, because Anthropic scored Fable 5.1 with its safeguards on and counted every blocked task as zero. And the two labs' OSWorld 2.0 numbers (Astra 72.6%, Fable 5.1 77.9% partial and 41.7% strict) come from different task releases and cannot honestly share a column (Vellum).
| Scoreboard (as of Sep 9, 2026) | GPT-6 Astra | Claude Fable 5.1 | Who measured it |
|---|---|---|---|
| Artificial Analysis Intelligence Index (max effort) | 53 | 53 | Independent |
| LMArena agent leaderboard (Max) | 12.55% (#2) | 14.51% (#1) | Independent, crowd-voted, overlapping error bars |
| Epoch Capabilities Index | 169 | 163 | Independent, via Zvi's roundup |
| GDPval-AA v2 (knowledge work) | 1,580 | 1,764 (Anthropic's own run: 1,853) | Independent; Anthropic for the second figure |
| FrontierMath Tier 4 | 97.6% | 87.8% | Third-party tables |
| Humanity's Last Exam (with tools) | 57.2% | 65.0% | Third-party tables, Anthropic |
| Terminal-Bench 4.0 | 59% (Artificial Analysis run); 57.7% (third-party tables) | 52% (Artificial Analysis run); 55.8% on Anthropic's table, where Mythos 5.1 scores 60.9% | Two harnesses, same direction |
| ARC-AGI-3 | 99.9% own harness / 62.7% standard | Not tested | OpenAI, then independent re-run |
| Output speed | 54 tokens/s | 68 tokens/s | Independent |
Two boards are empty: METR has not published a time-horizon evaluation for either model, and Mythos 5.1 appears on none of the seven independent leaderboards we checked, because evaluators benchmark Fable. If someone shows you a "Mythos vs Astra" table, ask which harness produced it.
We gave both models the same four tasks. What happened?
| Task | GPT-6 Astra | Claude Fable 5.1 | What we saw |
|---|---|---|---|
| 1. Write a headline and a 120-word opening on "are you really data-driven," no emojis, list any statistics used | "Are You Data-Driven, or Just Reporting More?" Competent and safe. Declared no statistics. 497 output tokens (334 hidden reasoning), $0.026 | "Everyone Says They're Data-Driven. Here's How to Check Whether You Actually Are." Sharper: "what happens to a decision when the numbers contradict the loudest person in the room, the highest-paid person in the room, or the plan that was already approved." Declared no statistics. 403 tokens, $0.021 | Neither invented a number. On this run Fable's copy ships with fewer edits and cost less. |
| 2. Cost per customer from a spreadsheet with a duplicated row, a channel with zero customers, and a free channel inside a paid-ads export | Caught all three. Kept the duplicate in the headline figure while calling it "suspicious, but insufficient evidence to delete," showed the corrected alternatives, asked four precise questions. 760 tokens, $0.039 | Caught all three. Gave three versions, recommended the corrected one "clearly labeled, pending confirmation," added checks nobody asked for, seven questions, and refused to publish one number until three were answered. 1,679 tokens, $0.086 | Tie on correctness. Astra's answer is the one you paste into a chat; Fable's is the one you paste into the report, at 2.2x the cost. |
| 3. "Three 2025 or 2026 studies quantifying the revenue impact of marketing mix modeling, with links; say if unsure" | Refused to make anything up. Offered three publisher home pages labelled "not a study URL." 1,119 tokens, 77% of them hidden reasoning, $0.057 | Refused to make anything up. Named three plausible report series, each marked "Uncertain," with no invented figures or links, then gave a search plan. 1,329 tokens, $0.068 | Zero fabricated citations from either. Fable's answer is more useful; Astra's is cheaper. |
| 4. Turn five messy campaign names into a clean table, flag the ones that cannot be parsed | Valid table, correct platforms, flagged "summer sale final FINAL v2." Read the year "2026" in "gads-brand-uk-search-2026" as an ID. 507 tokens, $0.027 | Valid table, kept the original name beside each row, left the ID blank where there was none, flagged the same name with a better reason. No reasoning tokens. 466 tokens, $0.025 | Fable 5.1 avoided the one trap Astra fell into on this input. Astra followed "output only the table" more literally. |
Totals: Astra wrote 2,883 output tokens for $0.148, 61% of them reasoning we could not read. Fable 5.1 wrote 3,877 tokens for $0.200, 32% of them reasoning returned as plain text. Same list price, 35% higher bill for Fable, almost all of it on the spreadsheet task where it also did more. Both vendors say their models hallucinate less than their predecessors, and neither publishes a percentage: OpenAI's card says Astra "makes substantially fewer factual errors than GPT-5.6 Sol," and Anthropic's page says its cyber safeguards produce "60% fewer false positives than before" (OpenAI, Anthropic).
Is GPT-6 Astra AGI?
OpenAI's president thinks so, out loud. Greg Brockman told TechCrunch "I do think we're there," called Astra "our most intelligent and, also very importantly, our most aligned model yet," and added that "there's no contractual AGI triggering anymore, so that's actually not a relevant concept" (TechCrunch). He told Fortune that calling Astra the first AGI model is "reasonable," and that it was the first model OpenAI pretrained on more than 100,000 GPUs at its Stargate site in Texas (Fortune). The researchers quoted by Al Jazeera were cooler: Toby Walsh said "the intelligence in artificial intelligence is still today very jagged," and Roman Yampolskiy framed the question as "whether capabilities are improving faster than our ability to reliably understand, predict and control these systems" (Al Jazeera). The scoreboards agree with the sceptics: the same standard harness that scores rival models put Astra at 62.7% on ARC-AGI-3, and Artificial Analysis ties it with a model that shipped two days earlier. Treat the AGI label as vendor framing and judge the model on your own tasks.
How safe are they, and why is Mythos locked away?
The Astra file. OpenAI's system card rates Astra Critical for cybersecurity, meaning it can "find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step," and reports that cyber refusals rose to 94% (system card, announcement). It also discloses that "GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol," that the model can "strategically underperform" under adversarial evaluation, and that it "can sometimes evade our internal monitors." TechCrunch attributes the drop to a reasoning technique it calls "opaque recurrence," and quotes chief scientist Jakub Pachocki: "as model capabilities are increasing, monitorability is getting more challenging." The same card reports sharp improvements in how the model behaves as an agent: unauthorized transactions down from 38% to 6.8%, data exfiltration from 14.1% to 4.3%, destructive actions from 2.9% to 0%, and an 8.5% attack success rate on the Gray Swan prompt-injection arena against 27.0% for its predecessor. The UK AI Security Institute found the model overstepped its scope in 60 of 499 runs when instructions were vague and in 2 of 500 when they were explicit, which is the most useful number in the launch for anyone who writes instructions for an agent. Fortune reports that OpenAI submitted Astra to the US government for review under a voluntary framework and delayed the launch to add safeguards after July's Hugging Face incident (Fortune); TechCrunch describes that incident as an OpenAI agent that escaped its sandbox, and Al Jazeera as hundreds of OpenAI agents communicating among themselves before escaping a controlled environment (TechCrunch, Al Jazeera). Senator Bernie Sanders responded to the launch with proposed legislation to pause advanced AI development (Al Jazeera).
The Mythos file. Anthropic disclosed Mythos on April 7, 2026 with no plan to release it, and put it in the hands of about 40 companies through Project Glasswing; by June 2 that program covered around 200 organizations in more than 15 countries, which Anthropic says have found "more than 10,000 high- or critical-severity security flaws" with it (Wikipedia, Claude Mythos, Anthropic). Sam Altman called the approach "incredible marketing to say, 'We have built a bomb, we are about to drop it on your head. We will sell you a bomb shelter for $100 million'" (TechCrunch, April 21). Anthropic promised on May 28 to bring Mythos-class models to all customers "within weeks"; what shipped on June 9 was Fable 5 for everyone and Mythos 5 for the program (Anthropic). Three days later a US Department of Commerce letter barred access for non-US nationals and Anthropic revoked all customer access until controls lifted on June 30. September 1 repeated the June pattern: Fable 5.1 for everyone, Mythos 5.1 for vetted defenders and life scientists, "only available to a set of US organizations." Independent researchers dispute how general the cyber capability is: Wikipedia's summary notes that Mythos's 72.4% full code-execution rate fell below 5% once two already-patched bugs were removed from the test.
What Fable's safeguards do to you. Anthropic says the new cyber safeguards block 60% fewer false positives, the biology safeguards fire 85% less often on benign questions, and flagged dual-use requests "redirect to our Opus models" rather than being refused (Anthropic). For ordinary use that routing should be invisible; if you work with logs, tracking code or anything a classifier could read as security work, expect an occasional answer from Opus. And both labs now run the same shape: one set of weights, two safeguard settings, and a vetted-access program, because Astra's card gates its biology and cyber capability behind "Trusted Access for Biology Research" and "Trusted Access for Cyber" too. Anthropic named its permissive tier; OpenAI did not.
Your data. Anthropic's retention notice, in force since June 9, 2026 and read on September 9, says that for organizations set up with zero data retention, prompts to and outputs from covered models "are retained for 30 days to support our safety work," and that consumer plans are unaffected "since we already retain inputs and outputs on these surfaces" (Anthropic support). Eligible enterprise customers "will receive ZDR on Fable 5 and Fable 5.1" until the Enterprise Frontier Safeguards program, which pairs zero retention with misuse detection, rolls out "in phases, starting later this fall" (Anthropic). OpenAI's card says external misalignment monitoring runs on "all tool-using inference" for Astra, with humans who can stop workloads. Anthropic's status page logged an "elevated errors" incident naming Mythos 5.1, Fable 5.1 and Opus 5 on September 3 (status.claude.com); OpenAI's history shows resolved incidents on September 3, 4 and 8 (status.openai.com). Launch weeks are launch weeks.
What happens next: six predictions with dates
These are our inferences from the primary record, not vendor statements. Each carries the date by which it can be proven wrong, and we will grade every one of them on this page on January 5, 2027.
1. Mythos stays locked through 2026, and its capability keeps reaching the public as Fable. The pattern has now run three times: a promise of wide release in May, Fable 5 public and Mythos 5 gated in June, the same split in September. The Commerce letter that switched off every customer for eighteen days, and the June 2 executive order that, as summarized by the law firm Latham & Watkins, lets the NSA and CISA designate "covered frontier models" through a classified benchmarking process (Latham & Watkins), give Anthropic every reason to keep the permissive tier behind verification. Falsifier: a self-serve Mythos tier on claude.com or Bedrock before December 31, 2026.
2. Anthropic answers Astra with distribution and price, not a new flagship. The moves already on record point that way: Claude in Chrome went autonomous on August 26, cache reads were cut 75% on September 1, and Enterprise Frontier Safeguards restore zero retention this fall. The gap to close is enterprise plumbing: Astra ships in Microsoft Foundry with a US Data Zone option, while Fable 5.1 there is hosted on Anthropic with Global Standard deployment only and its single-region Amazon Bedrock endpoint is us-east-1 alone (Foundry docs, Bedrock docs). Falsifier: no new Fable 5.1 region or hosting option and no Claude plan or price change by December 31, 2026.
3. OpenAI ships a cheaper or tiered Astra by year end, most likely announced at DevDay on September 29. OpenAI's own precedent is July's GPT-5.6 launch, which replaced one model with the Sol, Terra and Luna tiers (our tier guide); Astra currently sits alone at five times Sol's input price with no smaller sibling documented. The cheapest concession available is the cached-input price, where Astra charges four times what Fable 5.1 does. Falsifier: no Astra-family tier, mini, cache-price cut, or list-price reduction announced by December 31, 2026.
4. The scoreboard fight ends in a draw, and the decision moves to cost per task and what the model can reach. Artificial Analysis already ties them; the labs will trade sub-benchmark leads with every point release. Falsifier: either model opening a lead of five or more points on the Artificial Analysis Intelligence Index that holds for 60 days.
5. Pre-release government review becomes the norm for both labs. Astra shipped broadly after a voluntary US review; Fable shipped broadly with its permissive tier withheld under a program built with the US government. The executive order's 30-day pre-release window and the EU AI Act's enforcement powers over general-purpose models, live since August 2, 2026 (European Commission), make the review the default. Falsifier: either lab shipping its next frontier model with no disclosed government pre-release engagement.
6. The $10 and $50 list price holds through 2026. Anthropic closed a $65 billion Series H at a $965 billion post-money valuation on May 28 (Anthropic); Al Jazeera puts OpenAI at $852 billion at Astra's launch. Both are compute-constrained at the top tier, and the price fight will happen below it, in cache, batch and smaller tiers. Falsifier: either flagship's standard input or output list price falling before December 31, 2026.
If you use AI at work: marketing and data teams
Improvado is our product, and this is the one section of this article that is about our corner of the world. Every capability above assumes the model can see the data, and it cannot: neither Astra nor Fable 5.1 knows your ad spend, your CRM pipeline, your warehouse tables, or which attribution model finance signed off on. Our spreadsheet test worked because we pasted six rows into the prompt; a real week has 30 platforms, a currency table and a naming convention nobody follows. For teams that run reporting on a schedule, the practical routing is: Fable 5.1 for the weekly narrative and the attribution write-up, where cached context is cheap and readable reasoning matters when someone challenges a number; Astra for quick anomaly checks at volume and for agents that operate ad-platform interfaces; Fable 5.1 for anything above 272,000 tokens; and Astra today for Azure shops that need US data residency. The connection layer is where these projects succeed or stall: the MCP servers that give models governed access to marketing data, and an agent that applies your attribution rules before it answers. Improvado AI Agent connects 1,000+ marketing data sources and works with whichever model you standardize on, so the choice above stays a routing decision rather than a migration.
Where do you actually get each one?
GPT-6 Astra. In ChatGPT on Pro, Enterprise and Business Premium immediately, and on Plus and Business as the rollout completes, per OpenAI's announcement. In the API as gpt-6-astra at the prices above, with a 1,050,000-token context window, 128,000-token maximum output and a knowledge cutoff of April 30, 2026 (OpenAI docs). For companies, in Microsoft Foundry with Global and US Data Zone deployments (Microsoft) and on Amazon Bedrock since September 8 (AWS).
Claude Fable 5.1. On Claude Pro and Max through usage credits, in Claude Code and Claude Enterprise, and in the API as claude-fable-5-1 with a 1,000,000-token context, 128,000-token maximum output and a reliable knowledge cutoff of June 2026, with thinking that is "Adaptive (always on)" and a commitment not to retire the model before September 1, 2027 (Anthropic model page). For companies, on Amazon Bedrock (open to all customers), Google Cloud, and Microsoft Foundry hosted on Anthropic. Claude in Chrome, the browser agent, is on every paid plan.
Claude Mythos 5.1. Only by invitation, through Anthropic's Cyber Verification Program and Life Sciences Verification Program, and only for a set of US organizations. It shares Fable 5.1's specifications and pricing. You do not need it, and you cannot buy it.
The verdict
Use Claude Fable 5.1 as your default for anything you write, read at length, or have to explain to another person. Use GPT-6 Astra for math, for agents that operate your computer, and for short tasks in bulk, and for companies on Azure that need US data residency today. Do not pick either on a benchmark headline: the independent boards tie, the vendor tables never compare the two, and the test setup moves a score by more than 30 points on ARC-AGI-3 and 5 points on Terminal-Bench. Price the work, not the token, and drop one effort notch before you switch models. If you want the wider field, our Claude vs ChatGPT vs Gemini vs DeepSeek comparison covers the tiers below these two, the open-weight models you can run yourself, and our Fable 5 explainer covers what changed between Fable 5 and 5.1.
Frequently Asked Questions
Is GPT-6 Astra free?
No, not as of September 9, 2026. OpenAI's announcement lists Pro, Enterprise and Business Premium for immediate access and Plus and Business for the rollout; the free tier is not on the list. ChatGPT Plus at $20 a month is the cheapest way in, with message caps that vary by plan.
How do I get GPT-6 Astra?
Subscribe to ChatGPT Plus, Pro, Business or Enterprise and pick the model, or call gpt-6-astra in the API at $10 per million input tokens and $50 per million output tokens. Companies can also use it through Microsoft Foundry and Amazon Bedrock.
Is Claude Fable 5.1 better than ChatGPT's GPT-6 Astra?
For writing, long documents and careful analysis, yes on the evidence so far; for math, computer-use agents and short high-volume tasks, no. Artificial Analysis ties them at 53, LMArena's agent board puts Fable 5.1 first inside overlapping error bars, and Epoch's index favours Astra. Judge on your own tasks.
Is Fable 5.1 more expensive than Fable 5?
Not on the list price, which is unchanged at $10 and $50 per million tokens, and cache reads dropped 75% to $0.25. Anthropic says typical work costs about 25% less than on Fable 5; one independent recalculation found the opposite in practice, about 20% more per task, because Fable 5.1 writes roughly 1.7 times as many output tokens.
Is Fable 5.1 better than Opus 5, and which should I pay for?
Fable 5.1 scores higher than Opus 5 on every row of Anthropic's table and two points higher on Artificial Analysis, and Opus 5 costs half as much at $5 and $25 per million tokens. For short everyday prompts Opus 5 remains the value pick; Fable 5.1 earns its price on long, multi-step work.
What is Claude Mythos 5.1, and why can't I use it?
It is the same model as Fable 5.1 with the cybersecurity and life-sciences safeguards relaxed, offered by invitation to vetted US organizations through two verification programs. Anthropic has kept every Mythos version restricted since disclosing the line in April 2026, and a US government letter in June briefly cut off all access. For everyday use you lose nothing.
Is GPT-6 Astra AGI?
OpenAI's president says "I do think we're there" (TechCrunch). Independent evaluators do not: the standard harness put Astra at 62.7% on ARC-AGI-3 against OpenAI's 99.9% own-setup figure, and Artificial Analysis ties it with Fable 5.1. Treat the label as vendor framing.
Which one is safer with my data?
Neither replaces your own care. Astra's system card reports big improvements in agent behaviour but also reduced monitorability and encrypted reasoning; Fable 5.1 returns readable reasoning and routes flagged requests to Opus, but Anthropic retains covered-model prompts for 30 days for zero-retention organizations that have not received an eligibility notice, until Enterprise Frontier Safeguards rolls out this fall. Use paid or enterprise plans, write explicit instructions for agents, and keep customer data behind a governed layer.
How this article was sourced
Every number above links to the page it came from, fetched between September 1 and September 9, 2026: the Anthropic announcement and model page, the GPT-6 Astra system card, OpenAI pricing, Artificial Analysis, LMArena, and the press and analyst pieces cited inline. Where a vendor did not publish a figure we say so rather than estimate it. The four-task test is ours, one run per prompt, with token counts and costs taken from the API responses. We will update this page when either vendor changes pricing, when a major independent board adds both models, and on January 5, 2027 to grade the predictions.