Sources: Artificial Analysis (incl. Intelligence Index v4.3), DataCamp, BenchLM, OpenRouter, Roboflow Vision Evals, Vals AI, MindStudio, LLMLearner, Algoramming, CodingFleet, the-decoder, AA v4.3 announcement (LinkedIn) β fetched 2026-09-07, updated 2026-09-08
| Rate (per 1M tokens) | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|
| Input | $10.00 | $10.00 |
| Output | $50.00 | $50.00 |
| Cache read | $0.25 | $1.00 |
| Cache write (5 min) | $12.50 | $12.50 |
| Batch discount | 50% | 50% |
| Above 272K input | No surcharge | 2x input/cache, 1.5x output ($20/$2/$75) |
| Context window | 1,000,000 | 1,050,000 |
| Max output | 128K | 128K |
| Knowledge cutoff | June 2026 | April 30, 2026 |
| Release date | 2026-09-01 | 2026-09-03 |
| API model ID | claude-fable-5-1 | gpt-6-astra |
| Speed (tokens/sec) | 68.2 | 62.5 |
| Workload | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|
| Balanced: 1M in / 250K out | $22.50 | $22.50 |
| Retrieval sub-threshold: 10M in / 1M out | $150 | $150 |
| Retrieval over-threshold: 10M in / 1M out | $275 | $150 |
| Cache-heavy loop: 100K prefix x1000 reads | $201 | $126 |
| Effort | GPT-6 Astra (score/cost) | Claude Fable 5.1 (score/cost) |
|---|---|---|
| max | 61 / $1.67 | 66 / $3.76 |
| xhigh | 61 / $1.20 | 65 / $2.72 |
v4.3 update (2026-09-07): AA re-anchored the index. Fable 5.1 and Astra are now tied at 53. Cost per task: Astra $3.26 vs Fable 5.1 $7.63. See Β§7.5.
| Index | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|
| Coding Agent Index | 67 (Codex) | 70 (Claude Code) |
| Intelligence Index (max) | 61 | 66 |
| AA cost per Intelligence task (max) | $1.67 | $3.76 |
| Hallucination rate (max) | 51% | β |
| AA-Briefcase Elo gain | +80 | β |
| Category | Claude Fable 5.1 | GPT-6 Astra | Reading |
|---|---|---|---|
| Overall | 82.95 | 81.05 | overlap |
| Agentic | 78.7 | 70.4 | Fable leads |
| Knowledge | 86.4 | 81.7 | Fable leads |
| Coding | 84.2 | 75.3 | directional |
| Reasoning | 79.4 | 88.8 | directional |
| Benchmark | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|
| FrontierMath Tier 4 v2 | 97.6% | 87.8% |
| GPQA Diamond | 96.0% | 93.7% |
| Terminal-Bench Science 0.1 | 64.6% | 52.6% |
| Terminal-Bench 4.0 | 57.7% | 55.8% |
| DeepSWE v1.1 | 74.1% | 67.4% |
| FrontierCode 1.1 Main | 53.3% | 50.9% |
| HLE with tools | 57.2% | 65.0% |
| OSWorld 2.0 | 72.6% | 41.7% |
| ScreenSpot-Pro | 92.7% | 87.3% |
| AutomationBench | 41.4% | 31.4% |
| BenchCAD | 95.9% | 84.3% |
| ExploitBench | 100% | 70% |
| ARC-AGI-2 | 95% | 90% |
| CursorBench 3.2.0 | β | 73.4% |
| SWE-bench Pro | β | 81.2% |
| Toolathlon-Verified | β | 77.8% |
| LiveCodeBench (Vals) | β | 90.5% |
| BrowseComp | 91.5% | β |
| HealthBench Professional | 63.4% | β |
Choose GPT-6 Astra for: computer use, professional artifacts, math/science reasoning, cybersecurity defense, lower cost per task, Codex harness.
Choose Claude Fable 5.1 for: reasoning depth (HLE), cache-heavy agent loops, very large requests (no surcharge), Claude Code, highest independent reasoning score.
Source: Reddit r/better_claw Day-one numbers, one session, serving infra hours old. Also compares Claude Sonnet 5 ($3/$15).
Agent workloads are 80-95% cached context (same SOUL.md, schemas, history prefix every call) β effective input cost is the cache line, not the sticker.
| Input | Cached input | Output | |
|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 |
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 |
| Claude Sonnet 5 | $3.00 | $0.15 | $15.00 |
Test 1 β Tool calling under repetition (50 identical JSON-schema calls)
Test 2 β The "done" lie (6-step chain, step 4 guaranteed to fail)
Test 3 β Instruction survival past message 25 (tie across all three; nobody solves instruction decay)
Test 4 β Context honesty (~200K tokens, ask about something not in them)
Test 5 β Cost per real task (search, fetch 3 sources, synthesize, write summary)
Morning briefings, email triage, classification, drafting, research summaries, simple tool calling β everything a personal agent does 50x a day. On these, all three produce output indistinguishable in a blind read, and Sonnet does it at a quarter of the price.
Across 8 shared benchmarks, GPT-6 Astra leads on 5, Fable 5.1 on 3.
Domain means (0-100 scale, best published result each):
| Domain | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|
| Reasoning | 97.5 | 98.5 |
| Agent & tool use | 54.2 | 61.3 |
| Productivity | 31.4 | 41.4 |
Comparable shared benchmarks:
| Benchmark | Claude Fable 5.1 | GPT-6 Astra | Gap |
|---|---|---|---|
| ARC-AGI-1 | 97.5 | 98.5 | 1.0 Astra |
| ARC-AGI-2 | 90.0 | 95.0 | 5.0 Astra |
| Terminal-Bench 4.0 | 55.8 | 57.9 | 2.1 Astra |
| Terminal-Bench-Science 0.1 | 52.6 | 64.6 | 12.0 Astra |
| AutomationBench | 31.4 | 41.4 | 10.0 Astra |
Blended price at standard mix: both ~$20.00/1M (cache-read gap excluded from blend).
| GPT-6 Astra | Claude Fable 5.1 | Qwen 3.8 Max | |
|---|---|---|---|
| Input / Output (per 1M) | $10 / $50 | $10 / $50 | $2 / $6 |
| Cache read | $1.50 | $0.25 | $0.25 |
| Context | 1.05M | 1M | 1M |
| ARC-AGI-3 | 99.9% (stateful adapter) / 62.7% (stateless API) | β | β |
| OSWorld-Verified | β | 85.0% | 86.1% |
| DeepSWE | β | 80.0% | 56.6% |
| PaperBench | β | 88.8% | 93.0% |
| ExploitBench | 100% | β | β |
| CodeRabbit bug detection | +4% vs Sol, +22% vs Opus 5 | β | β |
Key caveat: Astra's ARC-AGI-3 score depends entirely on OpenAI's stateful Provider Adapter harness (99.9% with, 62.7% without) β the harness preserves reasoning state between API requests. Without it, the stateless API score drops sharply.
6 vision tasks, pooled at low effort. Astra wins overall but the two split the tasks 3-3.
| Overall | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|
| Vision Evals overall | 81.3% (#8 of 53) | 86.6% (#1 of 53) |
| Avg cost / sample | $0.035 | $0.030 |
| Avg speed / sample | 8.28s | 6.67s |
By task (low / high effort):
| Task | Claude Fable 5.1 | GPT-6 Astra | Winner |
|---|---|---|---|
| Object Detection | 61.4 / 65.0 | 82.1 / 83.6 | Astra (widest gap, 20.7 pts) |
| Counting | 69.4 / 73.0 | 80.2 / 81.1 | Astra |
| Identification | 97.9 / 96.9 | 89.6 / 92.7 | Fable 5.1 |
| OCR | 94.0 / 93.6 | 91.9 / 91.5 | Fable 5.1 |
| Data Extraction | 93.1 / 93.5 | 88.7 / 91.1 | Fable 5.1 |
| Reasoning | 72.0 / 73.1 | 87.2 / 91.2 | Astra |
Astra is both cheaper and faster per sample on the vision task mix.
Key independent analysis points:
ARC-AGI-3 harness dependency (confirmed):
Coding is a tie, not a takeover:
No broad intelligence jump:
FrontierMath caveat: 97.6% on Tier 4, but Epoch AI's harder 68-problem ErdΕs benchmark β Astra solved only 2 of 68 officially (rising to 5 with retries costing >$220,000 compute).
Cybersecurity jump: 100% ExploitBench, 42.4% ExploitGym, 88% SRE pass@1. First model to cross OpenAI's "Critical" threshold. Tradeoff: internal reasoning harder to monitor; Apollo Research flagged awareness-of-evaluation concerns.
(the-decoder / AA v4.2 coverage, 2026-09-05)
Superseded. v4.3 (Β§7.5) re-anchored the scale and now has Fable 5.1 and Astra tied at 53. The v4.2 figures below (Fable 5.1 57, Astra 55) are kept for historical context only. AA's live models page still shows these stale v4.2 numbers. Confirmed: Fable 5.1 tops, Astra 2nd (+4 pts on Sol).
v4.2 changes:
Where each shines:
| Sub-index | Leader | Detail |
|---|---|---|
| AA-Briefcase | Fable 5.1 & Opus 5 | Astra +85 Elo on Sol |
| GDP.pdf | GPT-6 Astra (33.2%) | Sol 28.2%, Fable 5.1 26.2% |
| Cost per task Pareto | 4 labs share (Anthropic, OpenAI, Meta, Z.AI) | β |
| Output token frontier | Astra dominates | most token-efficient near frontier |
Ranking: Fable 5.1 > Astra > Muse Spark 1.3 (Meta) > SpaceXAI > Kimi (Moonshot) > Z.AI > Google.
Source: AA official announcement (LinkedIn)
Major shift: Fable 5.1 and GPT-6 Astra are now TIED at 53.
| Intelligence Index v4.3 | Score |
|---|---|
| Claude Fable 5.1 (max w/ fallback) | 53 |
| GPT-6 Astra (max) | 53 |
| Claude Opus 5 (max) | 51 |
| Claude Fable 5 (w/ fallback) | 50 |
| Muse Spark 1.3 (max) | 48 |
| GPT-5.6 Sol (max) | 47 |
v4.3 changes (from v4.2):
v4.3 sub-benchmarks:
| Benchmark | GPT-6 Astra | Claude Fable 5.1 | Notes |
|---|---|---|---|
| Terminal-Bench 4.0 | 59.1% | 52.0% | vs Opus 5 49.0%, Sol 39.9% |
| AutomationBench-AA | 68.5% | β | vs Grok 4.6 66.7%, GLM-5.3 62.2% |
| Cost per task | $3.26 | $7.63 | Astra 57% cheaper |
Open weights: GLM-5.3 & Kimi K3 lead at 44 (9 pts behind the frontier pair); GLM-5.3 Flash 42, Qwen3.8 2.4T A95B 40, DeepSeek V4 Pro 36.
Cost Pareto frontier (v4.3): OpenAI occupies the majority β all five Astra effort levels offer the lowest cost per task at their intelligence level. Fable 5.1 (xhigh, max), GLM-5.3-Flash (42), and MiMo-V2.5-Pro (26) round out the frontier.
Note: AA's live models page still displays v4.2 figures (Fable 5.1 57, Astra 55) β that is cached/stale. The LinkedIn announcement is the authoritative v4.3 source.
| Claude Fable 5.1 | GPT-6 Astra | |
|---|---|---|
| Score | 74/80 (92.5%, #1) | 72/80 (90%, #3) |
| Token cost | ~$113 | ~$198 |
CodingFleet: Astra's launch demos operated KiCad, Unity, FreeCAD, and Blender directly. MindStudio: Astra connected to Blender via MCP, modeled an asset from scratch, exported it, imported into Unreal Engine 5 β handling the entire pipeline autonomously, verifying topology at each step.
Real user test (r/OpenAI, 2026-09-05): connected Astra to Blender via MCP, gave a prompt for a game asset, it created it from scratch. Astra's output notably better than Fable 5.1 on inanimate objects; humanoid/sculpted mesh struggled ("nowhere near the quality of the car"). Requires human final touchup on organic models.
| Test | Result |
|---|---|
| No Hype Assessment (YouTube), TB 4.0 max | Astra 56.7% vs Fable 55.8% β "basically the same"; cost $10.35 vs $19.50 |
| GDPval-AA v2 (Anthropic table) | Fable 5.1 185 vs Opus 5 182 vs Fable 5 172 vs Sol 171 |
| Output speed (CodingFleet) | Astra ~87 tok/s vs Fable ~67-69 β diverges from AA's 71.3 vs 70.5 "tie" |
| Fable 5.1 OSWorld caveat | Safeguards intervened on some tasks β scored zero, likely understating raw capability |
| Fable 5.1 front-end / code review | Consistently favored across Theo, MindStudio, KingBench |