Five Points That Became Two: What’s Wrong With The Astra Vs Fable Benchmark
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Five Points That Became Two: What’s Wrong With The Astra Vs Fable Benchmark on ThorstenMeyerAI.com

TL;DR

A recent review of the Astra versus Fable benchmark exposes significant inaccuracies caused by index revisions and architectural differences. The widely circulated comparison is based on outdated or misinterpreted data, leading to misleading conclusions about AI performance and economics.

Recent analysis reveals that the widely circulated comparison between GPT-6 Astra and Fable 5.1 AI models is based on outdated or misinterpreted benchmark data, leading to misleading conclusions about their relative performance and economics. The core issue stems from multiple revisions to the Artificial Analysis Index and architectural differences in how the models reason, which have not been properly accounted for in the comparisons.

The initial comparison suggested Astra lagged behind Fable 5.1 in the Artificial Analysis Intelligence Index by five points, with Astra supposedly scoring 61 and Fable 66. However, recent insights show that these figures are based on different versions of the index, which were updated around Astra’s launch, causing the scores to shift or become inconsistent. For example, the latest version of the index reports Astra scoring 55-57 and Fable scoring 54-57, indicating a much narrower gap that could be within a margin of error.

Further complicating the narrative, the original comparison conflated two different metrics. The AI community widely cited Astra as being economically superior due to lower token usage—$1.67 per task versus $3.69 for Fable—implying better efficiency. Yet, the Artificial Analysis report clarifies that Astra’s higher cost stems from its architecture, which reasons in latent space without producing chains of thought in tokens, making token counts an unreliable proxy for compute or intelligence. The efficiency gains are real but limited to specific coding tasks, not general intelligence.

Additionally, Astra’s architecture involves looping or recurrent reasoning that does not manifest as token output, meaning the token-based index measures are increasingly inaccurate. The models’ reasoning processes are internalized, and the token count no longer reflects the true computational effort. As a result, the comparison based solely on token usage is fundamentally flawed, and the narrative that Astra “attacks the economics” of intelligence is misleading outside narrow contexts.

At a glance
analysisWhen: developing; key revelations published i…
The developmentThe Astra vs Fable benchmark comparison has been shown to be flawed due to index revisions and architectural differences, causing confusion over AI performance and cost-efficiency.
Five Points That Became Two — Reality Check
AI Dispatch · Reality Check · 5 September 2026

Five points that became two: what’s wrong with the Astra vs Fable benchmark

The comparison everyone is quoting — Fable 66, Astra 61, “not a rounding error” — is built on numbers that were stale when written, measuring a quantity that no longer means what it used to, aggregated in a way that hides the reversals that matter. The benchmark isn’t broken. The way it’s being read is.

Problem 1 — the numbers moved: Index v4.1.1 → v4.2 during launch week
as quotedAA today (v4.2)
Claude Fable 5.16657 max effort
GPT-6 Astra6155 max · a third source says 60
gap: 5 2 “Five points is not a rounding error.” Two points, on an aggregate of ten evals that just swapped three of them (GPQA Diamond out; AA-Briefcase + GDP.pdf in), is exactly a rounding error. Not AA’s fault — revising an index is how you keep it honest. The error is downstream: quote the version, or don’t quote the number.
◆ Problem 3 — the tell: Astra’s effort dial isn’t connected to the engine
non-reasoning554.4M tok · $990
medium52
xhigh54$2,778
max55$3,020
Non-reasoning = max. Same score, 3× the cost. Because Astra is reported to be a looped / recurrent-depth transformer — it reasons in latent space, without emitting tokens. The Index prices cost in tokens, measures verbosity in tokens, computes time in tokens. For this architecture it’s counting the receipt, not the work. “140M vs 42M tokens” compares Fable’s verbalized reasoning to Astra’s post-loop output — an artefact, not an efficiency finding. Nobody outside OpenAI knows what the loops cost in GPU-seconds.
The other three problems
02
AA’s own conclusion is the opposite of the story
AA’s benchmarking note: Astra is 75% more expensive than GPT-5.6 Sol at max effort and “largely sits behind its predecessor on the Intelligence Index vs cost frontier.” Price went 2.5× ($4/$20 → $10/$50); token savings only partly offset it. The genuine efficiency win lives in one place: the Coding Agent Index, where Astra equals Fable 5 at under half the cost. “Astra attacks the economics” stretched a true coding result over an intelligence index where AA says the reverse.
04
“Max effort” isn’t the same experiment twice
Fable at max = more tokens. Astra at max = ~nothing (see ladder). And OpenAI’s docs say Astra does not support `none` reasoning effort — yet AA lists a “non-reasoning” score. The most efficient-looking config on the leaderboard may not be one you can buy.
05
The aggregate hides the reversals
Index: Fable +2. OpenAI’s own evals (self-reported): Astra ahead 6 of 7 — AutomationBench, BenchCAD, Terminal-Bench 4.0, DeepSWE, TB-Science, FrontierMath T4; Fable takes HLE+tools. A 6–1 task split became a two-point average, and the average became the story. Ten choices deep, two points is noise wearing a number.
✓ What actually changed — and it’s not on the leaderboard
Hallucination rate 92% → 51% on AA-Omniscience — a 41-point drop; matters more than any 2 Index points “Same headline price” hides cache read $1.00 vs $0.25 (4×) + a 25% cache-write premium — the line that dominates agentic bills Coding Agent Index: Astra = Fable 5 at < half the cost — real, and narrow
The take

Three things happened at once: the Index was revised (five became two), the architecture changed (tokens stopped being compute), and the aggregate did what aggregates do (6–1 became +2). A leaderboard position now tells you less than it ever has — and the more advanced the architecture, the less it tells you. Latent reasoning is only the first architecture to break the token proxy. So with your Astra access: ignore the Index number. Take your ten real tasks. Run both models at the effort setting you’ll actually pay for. Measure the bill including the cache line. Measure the failure rate — the 41-point hallucination drop is the one number here I’d bet money on. The benchmark can’t decide for you anymore.

Sources: Artificial Analysis Intelligence Index v4.2 / v4.1.1 model pages (Fable 5.1; Astra non-reasoning/low/medium/high/xhigh/max — scores, tokens, total cost, cost-per-task method incl. cache-write & reasoning tokens) and “Benchmarking GPT-6 Astra” (75% vs Sol, cost frontier, Coding Agent Index, 92%→51% hallucination, 2.5× price, cache terms); OpenAI GPT-6 Astra developer docs (`none` unsupported, cache-write billing, logprobs removed); Alan D. Thompson, The Memo 4 Sep 2026 (looped-transformer read, unconfirmed); OpenAI’s self-reported Astra-vs-Fable table; the circulating 66/61 comparison (pre-v4.2). Scores are version-dependent and were changing at time of writing. Not investment advice.
thorstenmeyerai.com

Implications for AI Performance and Benchmarking

This analysis underscores the importance of accurate benchmarking in AI development and the risks of relying on outdated or misinterpreted data. The flawed comparison has led to misconceptions about Astra’s efficiency and intelligence capabilities, potentially influencing investment, research priorities, and public perception. It highlights that architectural differences—such as reasoning in latent space versus tokenized reasoning—must be carefully considered when evaluating AI models.

For users and developers, understanding these nuances is critical for making informed decisions about model deployment and improvement. The case illustrates that raw token counts are insufficient as a measure of true AI efficiency, especially as models evolve architectures that process information differently. The broader lesson is that benchmarks must adapt to architectural innovations and be transparently versioned to avoid misleading conclusions.

Amazon

AI benchmarking analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Revisions and Architectural Shifts in AI Benchmarking

The Artificial Analysis Index has undergone multiple updates, including version 4.1.1 to 4.2, which involved removing some metrics like GPQA Diamond and adding others such as AA-Briefcase and GDP.pdf. Each update recalibrated the scores of models against a different evaluation basket, causing the absolute numbers to shift. This means that comparisons made using different versions are not directly comparable, and the initial five-point gap between Astra and Fable is no longer valid.

Moreover, Astra’s architecture is based on a looped transformer design, allowing it to reason internally without emitting tokens for some processes. This contrasts with models like Fable, which verbalize reasoning step-by-step, making token counts a more direct measure of computational effort. The shift in how Astra reasons fundamentally alters the meaning of token-based efficiency metrics, making previous comparisons obsolete.

Experts such as Sebastian Raschka and Alan Thompson have noted Astra’s architectural differences, emphasizing that the model’s internal reasoning mechanisms are not captured by traditional token-based benchmarks. This evolving understanding calls for more nuanced and architecture-aware evaluation methods.

“The comparison between Astra and Fable is based on outdated benchmark versions, and the numbers no longer reflect their true performance or efficiency.”

— Thorsten Meyer

Amazon

AI model performance evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Astra’s True Performance

It remains unclear how Astra’s internal reasoning process impacts real-world performance across diverse tasks, as current token-based benchmarks do not capture the full picture. OpenAI has not publicly disclosed detailed measurements of compute effort related to Astra’s latent loops, so the actual efficiency and cost-effectiveness are still uncertain. Additionally, the long-term implications of architectural differences on model scalability and general intelligence are still under investigation, with expert opinions diverging on how to best measure and compare these models.

Amazon

AI architecture comparison tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Benchmarking and Model Evaluation Strategies

Moving forward, the AI community is expected to develop more architecture-aware benchmarking methods that account for internal reasoning mechanisms like Astra’s latent loops. OpenAI and other organizations may publish detailed performance metrics that go beyond token counts, including GPU-seconds, energy consumption, and latency measurements. Further independent analyses are likely to scrutinize Astra’s real-world efficiency and capabilities, clarifying its position in the AI landscape. Meanwhile, the debate over how to fairly compare models with different architectures will continue.

Amazon

AI token usage efficiency calculator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why are the original Astra vs Fable benchmark numbers misleading?

The numbers are based on different versions of the Artificial Analysis Index, which were updated around Astra’s launch, causing the scores to shift. Additionally, the comparison conflates token-based efficiency with architectural differences, making the original figures inaccurate for assessing true performance or cost-effectiveness.

What is the main reason Astra’s token count does not reflect its actual compute effort?

Astra’s architecture reasons internally in latent space and uses looping mechanisms that do not produce tokens during reasoning, so token counts do not accurately measure the model’s computational effort or intelligence.

Will future benchmarks better capture Astra’s architecture?

Yes, experts expect new evaluation methods that incorporate hardware metrics like GPU-seconds and energy use, which will better reflect Astra’s internal reasoning processes and overall efficiency.

How does this analysis affect the perception of Astra’s capabilities?

It suggests that Astra may be more cost-efficient in specific coding tasks but does not necessarily outperform other models in general intelligence, especially when using token-based metrics that are no longer fully applicable.

What should researchers focus on to improve AI benchmarking?

They should develop metrics that account for architectural differences, internal reasoning mechanisms, and real compute effort, moving beyond token counts to achieve fairer, more accurate comparisons.

Source: ThorstenMeyerAI.com

You May Also Like

Show HN: Distilling DeepSeek into GPT-OSS doesn’t transfer censorship. Try it

A recent demonstration shows that distilling DeepSeek into GPT-OSS-120B does not carry over censorship features, raising questions about model control and safety.

Inkling: Our Open-Weights Model

AI researchers have announced ‘Inkling,’ an open-weights model designed for transparency and customization, marking a significant step in AI development.

Prevent Cognitive Debt By Manually Retyping LLM-generated Code

Experts recommend manually retyping AI-generated code to prevent errors and cognitive overload, improving developer accuracy and software quality.

GPT‑Live

OpenAI announces GPT‑Live, a new real-time chat interface enabling instant AI interactions, with beta testing now open to select users.