AIThis post was created with the assistance of artificial intelligence (AI).

By Thorsten Meyer

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

For two weeks, Alibaba’s biggest-ever model existed as a slogan.

“Second only to Fable 5.” A 2.4-trillion-parameter figure. A paid preview endpoint. No benchmark table, no model card, no licence, no active-parameter count.

Today, 3 August, the slogan grew a spec sheet. Alibaba made Qwen3.8-Max broadly available, published the full benchmark table it had withheld since the July preview, and confirmed that open weights ship next week — alongside a second checkpoint, Qwen3.8-27B, that matters more for anyone running their own hardware than the flagship does.

The numbers are genuinely strong. They are also genuinely selective. Both halves of that sentence deserve equal weight.

AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Two weeks of theatre, reconstructed

The launch choreography is worth recording, because it was choreography.

On 17 July, Moonshot released Kimi K3 — 2.8 trillion parameters, the model that briefly rattled US tech stocks and then suspended new subscriptions when demand outran capacity. On 18 July, an anonymous model called "kaleb" appeared on the Code Arena leaderboard, introducing itself as "Claude" — a distillation artifact — until the community identified it within a day by a quirk unique to Alibaba's Qwen tokenizer. On 19 July, during the World AI Conference in Shanghai, Alibaba confirmed it: kaleb was Qwen3.8-Max, previewed in stealth.

The Sunday announcement carried the "second only to Fable 5" claim and nothing to check it against. Alibaba's shares rose as much as 5.4 percent on the Monday anyway. The preview was purchasable through the Token Plan at ten percent of standard pricing — a real endpoint with real limits (a 983,616-token context window, 131,072 maximum output, thinking always on with three effort settings) and no paper trail.

That was the state of play until this morning. Announcing the claim on a Sunday and the evidence fifteen days later is a strategy, and it worked: two weeks of coverage ran on Alibaba's framing because there was nothing else to run on.

Amazon

AI language model development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the spec sheet actually says

The confirmed shape: 2.4 trillion total parameters, roughly 95 billion active per query, sparse mixture-of-experts, built on the Qwen3.5 architecture. Multimodal in — text, image, video — text out. The active-parameter figure is the number nobody had, and it settles the most important open question: this is a roughly-95B-compute model wearing a 2.4T coat, about 4 percent of the network firing per token.

The benchmark table, on Alibaba's own harness, splits cleanly into three stories.

Where it leads: Terminal-Bench 2.1 at 86.6 — ahead of both Claude Opus 4.8 and Claude Fable 5 at 84.6, behind only GPT-5.6 Sol at maximum effort at 88.8. PaperBench at 93.0, top of the table. IFBench at 82.8. The multimodal and agentic rows are where the model shines: OSWorld-Verified at 86.1, Parametric CAD Bench at 91.5, OmniDocBench 1.5 at 92.1. Alibaba also demonstrated the long-horizon pitch directly: the model reproduced all six main results of a research paper, then tested eighteen of its own ideas across four rounds and beat the paper's method on AIME24 by 2.7 points.

Where it trails, badly: SWE-bench Pro at 67.7 against Fable 5's 80.0. FrontierSWE at 73.5 against Fable 5's 88.8. These are not rounding errors — they are twelve- and fifteen-point gaps on exactly the deep software-engineering benchmarks that the "second only to" framing invites you to assume are covered.

Where it barely moved: GPQA Diamond at 92.6, up from Qwen3.7-Max's 92.4. The reasoning ceiling did not lift. What lifted — enormously — is agentic execution against its own predecessor: DeepSWE from 21.6 to 56.6, FrontierSWE from 40.7 to 73.5, JobBench from 31.3 to 53.4.

So the honest one-line read: "second only to Fable 5" is true on the rows Alibaba chose and false on the rows it didn't. This is not scandalous — every lab's launch table does this — but it means the claim that moved a share price is a claim about a subset. The clearest genuine achievement is different and less quotable: Alibaba took its own model from unusable to competitive on long-horizon agent work in one generation, mostly through RL-environment scaling, and published the trajectory.

Amazon

multimodal AI model hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The lane question

For this desk, the open-weight status is the operative fact, and it is now three-tiered.

The hosted API is live today, OpenAI- and DashScope-compatible — a base-URL change for anyone who wants to A/B it.

The 2.4T weights, due next week, are a gesture more than a deployment option. Whatever the licence turns out to be — and it is still unpublished, which matters, given that Qwen's open models have historically shipped Apache 2.0 while Kimi K3's licence carries revenue and attribution triggers — a 2.4-trillion-parameter checkpoint is a multi-node datacenter artifact. With the activated count at 95B, even aggressive quantisation puts the memory footprint far beyond any single machine. Nobody reading this self-hosts it. Its release is real openness and it is also a flag planted: the largest open-weight model ever shipped, if it ships.

Qwen3.8-27B is the row that belongs in the local-first lane. A 27B checkpoint distilled from whatever made the flagship's agentic numbers jump is precisely the shape of model that runs on a single high-memory machine as a daily driver — and precisely the class where this desk's inference actually happens. Whether the agentic gains survive the compression is the question that decides whether next week matters. There is no benchmark table for the 27B yet. There is a pattern, though: the interesting Qwen releases have consistently been the mid-size ones, and the flagship's job has been to make headlines that the 30B-class models then monetise in deployments.

Amazon

large parameter AI model weights

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Bull and bear

The bull case: the agentic-generation jump is real and internally consistent across a dozen rows; the RL-environment scaling story gives a mechanism, not just a number. The active-parameter disclosure and full table — even fifteen days late — is more than Kimi K3 shipped at launch. If the 2.4T weights land under Apache 2.0, the ceiling of what "open weight" means moves permanently, and every argument that frontier capability requires closed weights loses another data point. And the 27B sibling could be the best local agent model available, on hardware people already own.

The bear case: every number above is Alibaba's own harness, and agent benchmarks are the most harness-sensitive numbers in the field — independent testing of Kimi K3 already tempered that model's launch claims substantially, and there is no reason to expect this table survives contact with third parties intact. The deep-SWE gaps mean the flagship coding-agent use case — the one that pays — still belongs to Fable 5 by a wide margin. "Next week" is a promise from a company that spent two weeks not publishing a benchmark table it possessed. And the licence is still unwritten: until it exists, "going open-weight" is a press strategy, not a property of the model.

There is also a quieter structural point. Kimi K3, DeepSeek's 0731, and now Qwen3.8-Max arrived within seventeen days of each other, each timed against the others, each claiming adjacency to Fable 5 — a model the US temporarily placed under export controls, which has made "close to Fable 5" the season's universal marketing coordinate. When three labs in three weeks all measure themselves against the same restricted model, the benchmark being run is partly geopolitical. That contest is real, but it is not the same thing as your workload.

Amazon

AI benchmark testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What I'll actually do

Nothing, until three things exist: the licence text, the 27B benchmark table, and one independent harness run of the flagship's agentic rows. The API is a base-URL change whenever a comparison is worth running; the 2.4T weights are not a deployment option for this fleet regardless of licence; and the 27B — the release that could genuinely change the daily-driver calculus here — is currently a name in a press release.

The preview phase of this launch ran on a slogan for fifteen days because nothing checkable existed. The GA phase should not get the same courtesy. The table is out; now the table gets tested.


Sources: Alibaba Qwen team release and benchmark table via MarkTechPost, 3 August 2026; The Decoder on the 95B active-parameter figure, weights timeline, and RL-scaling account, 3 August 2026; Bloomberg on the release and Fable 5 comparison, 3 August 2026; MarkTechPost and eesel coverage of the 19 July WAIC preview; Wan27 and community reporting on the "kaleb" stealth identification; contemporaneous coverage of Kimi K3's launch, licence terms, and independent evaluations. All Qwen3.8-Max performance figures are Alibaba's own harness; the open-weight licence was unpublished at time of writing. Point-in-time as of 3 August 2026. Not purchasing advice.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Superintelligence for everyone – research overview (July 31 2025)

AIThis post was created with the assistance of artificial intelligence (AI). Buying…

Three Shots on Goal: The Warning Shot We Almost Didn’t Get

AIThis post was created with the assistance of artificial intelligence (AI).The Hugging…

Mistral Forge: Owning the Model, Not Just Renting the API

AIThis post was created with the assistance of artificial intelligence (AI).For two…

Startup Sofa Briefing: The Real Cost of Starting Up (It’s Not Just Money)

AIThis post was created with the assistance of artificial intelligence (AI). Buying…