In 30 days, Anthropic, OpenAI, and Google DeepMind shipped four frontier models. Claude Fable 5.1 arrived on September 1. GPT-6 Astra followed on September 3. GPT-6.1 Sol and Gemini 4 Argon landed at the month’s end. On September 28, OpenAI canceled GPT-6.1 Astra after internal scope and authorization tests. GPT-6 Astra remains OpenAI’s top model for now. We place the models side by side: benchmark scores overlap more than launch posts suggest. But pricing, access rules, and cost per task diverge.
What changed in the lineup this September?
The inflection was OpenAI canceling GPT-6.1 Astra on September 28. GPT-6 Astra stays the company’s flagship. The other three releases remain available, so the menu widened while gaps narrowed.
Timeline is simple: Claude Fable 5.1 on September 1, GPT-6 Astra on September 3. GPT-6.1 Sol and Gemini 4 Argon arrived in the last days of September. We covered each launch alone; now we compare them shoulder to shoulder.
Benchmarks sit closer than launch narratives imply. Leaders change by task: long-horizon coding, knowledge work, computer use, vulnerabilities, or terminal agents.
Differences come from prices, limits, and access. Lists look similar, but real task cost and cache reads reshape the order. Add access constraints and the front-runner shifts. Ready to dive into the numbers?
Specs, pricing, and access: where do they split?
Base rates match for Astra and Fable 5.1: $10 input and $50 output per 1M tokens. Sol and Argon price at one-fifth of that: $2/$10. Argon’s price is introductory and doubles later.
Cached input matters most for agents. Astra charges $1.00 per 1M cache reads. Fable 5.1 is $0.25. Sol is $0.10. Argon’s intro rate is $0.10. In long agent loops, these gaps compound fast.
Limits and context windows are close: 1.05M in Astra and Sol, 1M in Fable 5.1. Max output per response is 128K across three models, with one outlier: Argon allows 1M. That is the only structural difference for very long replies.
Access also diverges. Argon is available only through the Fairwind program. Astra’s full offensive security capability sits behind OpenAI’s Daybreak; the public release refuses advanced offensive tasks. Anthropic gates its unrestricted twin, Claude Mythos 5.1, behind trusted access programs.
Agents resend system prompts, tool schemas and history on every step. Astra’s $1.00 cache read is 4x Fable 5.1’s and 10x Sol’s.
Who actually leads on benchmarks?
No model sweeps the board. Argon leads knowledge work and long-horizon software engineering. Astra leads frontier software engineering and computer use.
Numbers read as follows: DeepSWE v1.1 — Argon 77.9%, Astra 74.1%, Fable 5.1 at 67.4%. Vals Index (finance, coding, legal, and tax work): Argon 68.9%, Fable 5.1 at 65.8%, Astra 63.1%. FrontierSWE v2: Astra in front at 65.5%, with Argon at 55.0% and Fable 5.1 at 56.3%.
Terminal-Bench 4.0: Astra 58.2%, Fable 5.1 at 57.9%, Argon at 57.4%. CWE-bench v1: tie at 68% for Argon and Astra, with Fable 5.1 at 58%. OSWorld-2.0: Astra at 72.6%, Argon at 69.2%; Google’s table does not list Fable 5.1.
GPT-6.1 Sol does not appear in Google’s table, but OpenAI’s numbers place it close to Astra. On DeepSWE v1.1, Sol matches Astra at roughly one-fifth the cost. On the OSWorld 2.0 offline set, Sol lands within 2.1 points at about one-seventh the task cost. On AutomationBench 1.0.6, Sol scores 2.2 points above Claude Opus 5.5 at medium effort. On Terminal-Bench Science 0.1, Astra still leads at 68.1%, and OpenAI recommends it for the hardest research. Independent signals differ on raw intelligence: on the Intelligence Index, Astra scores 61 while Fable 5.1 sits 5 points higher; on the coding-agent index in Claude Code, Fable 5.1 scores 70 against Astra’s 67. Artificial Analysis also reports Argon equals Astra on the Intelligence Index. On ARC-AGI-2, Astra scores 95% while Fable 5.1 scores 90%.
Cost per task: same list price, different bill
Artificial Analysis measures Claude Fable 5.1 at $9.18 per task versus $4.72 for GPT-6 Astra. That is about 1.9x despite equal list prices. The gap comes from token volume, not rates.
When prices are equal, a higher per-task bill means more tokens burned. Anthropic also notes its newer tokenizer yields roughly 30% more tokens for the same text.
Cheaper models shift the picture further. Artificial Analysis reports Argon equals Astra’s Intelligence Index at 60% of Astra’s task cost using intro prices. On Terminal-Bench Science, OpenAI reports $5.47 per task for Sol against $23.80 for Astra.
Caching can flip rankings for agentic loops. Consider a 200K-token cached context before outputs: Astra at $1.00/1M costs $0.20 per step and $20 per 100 steps; Fable 5.1 at $0.25/1M is $0.05 and $5; Sol and Argon (intro $0.10/1M) are $0.02 and $2. Astra’s 200K context stays under its 272K long-prompt threshold.
Caching can reverse the ranking for agents. Measure both on your own traces before you pick on per-task headlines.
How to choose: leaders by workload
Pick by workload, not by leaderboard rank. GPT-6.1 Sol is the default for most teams: lowest public price with Astra-like performance on key sets.
Hardest open-ended engineering? GPT-6 Astra leads FrontierSWE v2 at 65.5%. Computer use and browser agents? Again Astra: OSWorld-2.0 at 72.6%, with Sol within 2.1 points. Very long single outputs for refactors and full reports? Only Gemini 4 Argon offers 1M output tokens per response.
High-volume coding agents, CI bots, and PR review? GPT-6.1 Sol: it matches Astra on DeepSWE v1.1 at $2/$10 with $0.10 cache reads. For legal, finance, and business automation, pick Gemini 4 Argon: it leads Vals Index (68.9%), AutomationBench (51.3%), and Harvey Legal (19.6%). For vulnerability finding and patching, choose Gemini 4 Argon or GPT-6 Astra, tied at 68% on CWE-bench v1.
Most numbers here are vendor-reported, and vendors disagree at the margins.
Two rows hinge on access: today, Argon is available only to Fairwind cyber defenders. Astra’s full offensive-security capability sits behind Daybreak; the public build refuses advanced offensive cyber tasks. Anthropic also gates Claude Mythos 5.1 behind trusted programs. Before switching, check the fine print: Anthropic ran Fable 5.1 with production safeguards, which likely lowered OSWorld and AutomationBench; Argon’s $2/$10 is temporary and moves to $4/$20; Argon’s context window is undisclosed; per-task costs are a snapshot and shift with effort. Anthropic also has a cheaper option: our Argon coverage cites reports that Claude Opus 5.5 beats Fable 5.1 on key agentic benchmarks at a lower API price and leads Terminal-Bench 4.0 at 66.4%.
Based on source material.