Second Only To Fable 5: Qwen3.8-Max Finally Shows Its Numbers — And The Claim Gets Complicated

📊 Full opportunity report: Second Only To Fable 5: Qwen3.8-Max Finally Shows Its Numbers — And The Claim Gets Complicated on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba’s Qwen3.8-Max model has disclosed its benchmark scores, confirming a 2.4 trillion-parameter size and strong performance in several tests. The open weights are set to be released next week, marking a significant step in large-language model availability.

Alibaba has publicly released the benchmark scores for its Qwen3.8-Max model, confirming it as the second-largest model after Fable 5. The company detailed the model’s specifications, including 2.4 trillion total parameters and a 95-billion active-parameter count per query, and announced the upcoming release of open weights next week. This marks a significant milestone in large-language model development and availability.

On 3 August, Alibaba published the full benchmark table for Qwen3.8-Max, revealing it as a model with 2.4 trillion total parameters and a roughly 95 billion active-parameter count. The model is built on the Qwen3.5 architecture, employing sparse mixture-of-experts techniques, and supports multimodal input — text, images, and videos — with text output.

The benchmark results place Qwen3.8-Max as second only to GPT-5.6 Sol at 88.8 in the Terminal-Bench 2.1, outperforming Claude models but trailing GPT-5.6. It also scored highly in other tests, such as PaperBench at 93.0 and OSWorld-Verified at 86.1. Alibaba demonstrated the model’s capabilities by reproducing key research results and outperforming some of its own previous benchmarks, especially in long-horizon agentic tasks.

While the model’s active parameters and benchmark scores are confirmed, the open weights are scheduled for release next week. The weights will be a 2.4 trillion-parameter checkpoint, which is a multi-node data center artifact, though a smaller 27B version, Qwen3.8-27B, will be available for local deployment on high-memory machines. The licensing details remain unpublished, but historically, Alibaba’s open models have used Apache 2.0 licenses.

At a glance
updateWhen: announced August 3, 2023; benchmark dat…
The developmentAlibaba announced the full benchmark results for Qwen3.8-Max, confirming its 2.4 trillion parameters and competitive performance, with open weights arriving soon.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Qwen3.8-Max's Benchmark Results

The release of benchmark scores and upcoming open weights signals Alibaba's strategic move to compete in the large-language model space, especially against models like Fable 5 and GPT-5.6. The detailed performance data confirms the model's capabilities in multiple AI benchmarks, emphasizing its potential for enterprise and research applications.

Particularly notable is the model's demonstrated improvement in agentic tasks, where it shows significant progress over previous versions. The upcoming open weights could enable wider adoption and experimentation, influencing the development landscape and competitive dynamics among AI providers.

Amazon

high-memory GPU for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba's Large-Language Model Development

Over the past two weeks, Alibaba's Qwen3.8-Max was shrouded in secrecy, initially introduced through a stealth preview and a slogan claiming it was “second only to Fable 5,” a 2.8 trillion-parameter model. The model was identified as Qwen3.8-Max during the World AI Conference in Shanghai on 19 July when Alibaba confirmed its existence. Prior to this, the model appeared anonymously on the Code Arena leaderboard as “kaleb,” with community members quickly recognizing its unique tokenizer signature.

Alibaba's strategy involved a staged reveal, culminating in the full benchmark publication today. The model's architecture, based on Qwen3.5, employs sparse mixture-of-experts techniques, and the company has emphasized its multimodal capabilities and agentic performance. The upcoming release of open weights marks a significant milestone in making such large models accessible for broader use.

"We are committed to open AI development and will release the open weights next week, enabling broader experimentation and deployment."

— Alibaba spokesperson

Amazon

large language model open weights

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Licensing and Deployment

Details about the licensing terms for the 2.4 trillion-parameter open weights are still unpublished, raising questions about usage rights and restrictions. It is also unclear whether the open weights will include the full model or a compressed version suitable for local deployment. The performance of the 27B checkpoint on agentic and deep software benchmarks remains to be seen, as no benchmark data has been released for it yet.

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s Model Release and Industry Impact

Alibaba is scheduled to release the open weights of Qwen3.8-Max next week, which will allow researchers and developers to evaluate its capabilities firsthand. The company may also publish additional benchmark data for the 27B version, clarifying its performance in local deployment scenarios. Industry analysts will closely watch how the open release influences AI development, competition, and adoption across sectors.

Amazon

multimodal AI development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled for release next week, with exact timing to be confirmed.

How does Qwen3.8-Max compare to other large models like GPT-5.6 or Fable 5?

In benchmark tests, Qwen3.8-Max ranks second to GPT-5.6 Sol in Terminal-Bench 2.1 and outperforms Fable 5 in several metrics, though it trails in some software engineering benchmarks.

Will the open weights be fully open-source?

Details about licensing are still unpublished, but historically Alibaba's open models use licenses like Apache 2.0, though the upcoming release's licensing terms remain to be confirmed.

What is the significance of the 27B checkpoint?

The 27B version is designed for local deployment on high-memory machines, offering a more accessible option for practical use, though its performance benchmarks are not yet available.

What does the performance of Qwen3.8-Max imply for AI development?

The model's strong benchmark scores and agentic improvements suggest it could influence future AI research and deployment, especially if open weights enable broader experimentation.

Source: ThorstenMeyerAI.com

You May Also Like

The Financial Edge: Choosing Between Forge And Self-Hosting For Sovereign AI

Analysis of the costs and capabilities of Forge’s managed sovereignty platform versus self-hosted AI models in 2026.

15 Best AI Tools For Automating Workflows In 2026

Discover the 15 best AI tools for automating workflows in 2026, including no-code and developer-focused options, with insights on their applications and significance.

South Korea to invest $576 billion in AI chip production with Samsung and SK Hynix

South Korea plans a $576 billion investment in AI chip production, involving Samsung and SK Hynix, to strengthen its semiconductor industry and global competitiveness.

The Menu: What Ten Answers Reveal

A detailed analysis of ten jurisdictions’ responses to automation and AI, revealing diverse approaches to income, capital, work, skills, and institutions.