Second Only To Fable 5: Qwen3.8-Max Finally Shows Its Numbers — And The Claim Gets Complicated

📊 Full opportunity report: Second Only To Fable 5: Qwen3.8-Max Finally Shows Its Numbers — And The Claim Gets Complicated on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, revealing detailed benchmark results that position it as the second-best model after Fable 5. The model boasts 2.4 trillion parameters, with strong performance in multimodal and agentic tasks, though some limitations remain.

Alibaba officially released the full benchmark table for Qwen3.8-Max on August 3, confirming it has 2.4 trillion parameters and ranks as the second-highest model after Fable 5, based on their own tests. The company also announced that open weights for the model will be available next week, marking a significant step in AI model transparency and accessibility.

Following two weeks of speculation and partial previews, Alibaba published the detailed benchmark results for Qwen3.8-Max, revealing a model built on the Qwen3.5 architecture with a sparse mixture-of-experts design. The model features approximately 95 billion active parameters per query, operating within a 983,616-token context window, and supports multimodal input including text, images, and video.

The benchmark table shows Qwen3.8-Max achieving top scores on several tests, including Terminal-Bench 2.1 at 86.6, surpassing Claude Fable 5 and only behind GPT-5.6 Sol at 88.8. It also leads in PaperBench at 93.0 and performs well in multimodal and agentic tasks, with notable scores in OSWorld-Verified and Parametric CAD Bench. However, it trails significantly in deep software-engineering benchmarks like SWE-bench Pro and FrontierSWE, with gaps of 12-15 points compared to Fable 5.

The company emphasized that the 2.4 trillion parameters are sparsely activated, with the active parameters per query around 95 billion, indicating a model that combines large-scale capacity with efficiency. They also demonstrated the model’s ability to reproduce research results and outperform previous versions on long-horizon agent tasks, marking a major improvement over prior iterations.

At a glance
updateWhen: announced August 3, 2023; benchmarks re…
The developmentAlibaba has released detailed benchmark results for Qwen3.8-Max, confirming its parameter count and performance, and announced open weights will be available next week.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Results for AI Leadership

Alibaba’s detailed benchmark release confirms its position as a leading AI developer, with a model that rivals GPT-5 in several areas. The disclosure of the 2.4 trillion parameters and performance metrics signals a shift toward more transparent and competitive large-scale models. The availability of open weights next week will enable broader access and testing, potentially impacting AI deployment and innovation across industries. However, the model still shows limitations in software engineering tasks, highlighting ongoing challenges in specialized AI capabilities.

Amazon

AI model benchmark tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Alibaba’s Model Development and Benchmarking Strategy

Over the past two weeks, Alibaba’s Qwen3.8-Max was shrouded in secrecy, initially revealed through a stealth preview and an anonymous model called “kaleb” on the Code Arena leaderboard. The model was confirmed during the World AI Conference in Shanghai on July 19, after a series of teasers and leaks. The company’s approach involved staged announcements, culminating in the full benchmark release on August 3, which included detailed performance data for the first time.

This follows a recent surge of large model launches, including Moonshot’s Kimi K3 with 2.8 trillion parameters, and reflects Alibaba’s strategic emphasis on multimodal and agentic capabilities. The model’s design leverages sparse mixture-of-experts architecture, aiming to balance size, efficiency, and performance, especially in long-horizon reasoning tasks.

"Qwen3.8-Max demonstrates our commitment to advancing multimodal and agentic AI capabilities, with open weights coming next week."

— Alibaba spokesperson

Amazon

multimodal AI development kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Licensing and Real-World Use

It is still unclear what the licensing terms for the open weights will be, as Alibaba has not yet published the license details. The actual performance of the 27B checkpoint in real-world deployment remains to be tested, especially regarding whether agentic gains are preserved after compression. Additionally, the long-term impact of the model’s limitations in software engineering benchmarks is uncertain.

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps: Open Weights Release and Broader Testing

Alibaba plans to release the open weights of Qwen3.8-Max next week, enabling researchers and developers to evaluate its capabilities firsthand. The upcoming release will likely trigger a wave of independent benchmarking and application development. Further updates on licensing, deployment, and performance in diverse tasks are expected in the coming months as the model is integrated into broader AI ecosystems.

Amazon

AI model performance testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the key specifications of Qwen3.8-Max?

Qwen3.8-Max has 2.4 trillion total parameters with approximately 95 billion active parameters per query, built on a sparse mixture-of-experts architecture, supporting multimodal input and long context windows.

When will the open weights be available?

Alibaba announced that open weights for Qwen3.8-Max will be shipped next week, enabling broader access for research and deployment.

How does Qwen3.8-Max compare to other models like GPT-5 or Fable 5?

According to Alibaba’s benchmarks, Qwen3.8-Max ranks second after GPT-5.6 Sol on Terminal-Bench, and surpasses Fable 5 in several multimodal and agentic tasks, though it trails in deep software engineering benchmarks.

What are the main limitations of Qwen3.8-Max?

The model underperforms significantly on deep software-engineering benchmarks and its agentic capabilities may be affected by compression in the 27B checkpoint. Its licensing terms remain undisclosed.

Why is the benchmark release significant?

The detailed benchmark results validate Alibaba’s claims of high performance and transparency, potentially reshaping competitive dynamics among large AI models.

Source: ThorstenMeyerAI.com

You May Also Like
Five Levers, Many Hands

Five Levers, Many Hands

Analysis of how countries respond to AI-driven labor shifts using five key tools, highlighting differences and uncertainties in the global response.
The Model Is Only 10%: The Real Lesson of the New SDLC

The Model Is Only 10%: The Real Lesson of the New SDLC

A new Google whitepaper emphasizes that AI models account for only 10% of system behavior; verification and configuration are the real focus.
RHEO On The Web: Find Your Flow

RHEO On The Web: Find Your Flow

Discover RHEO’s web version — a private, instant, browser-based fluid simulation for calm, creativity, and ambient displays, with no sign-up or downloads.
Forezai · TradingAgents: A Trading Firm Made of Agents

Forezai · TradingAgents: A Trading Firm Made of Agents

Thorsten Meyer AI released TradingAgents, an Apache-2.0 research framework that models an AI trading desk with debate and risk vetoes.