By Thorsten Meyer
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
For two weeks, Alibaba’s biggest-ever model existed as a slogan.
“Second only to Fable 5.” A 2.4-trillion-parameter figure. A paid preview endpoint. No benchmark table, no model card, no licence, no active-parameter count.
Today, 3 August, the slogan grew a spec sheet. Alibaba made Qwen3.8-Max broadly available, published the full benchmark table it had withheld since the July preview, and confirmed that open weights ship next week — alongside a second checkpoint, Qwen3.8-27B, that matters more for anyone running their own hardware than the flagship does.
The numbers are genuinely strong. They are also genuinely selective. Both halves of that sentence deserve equal weight.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Two weeks of theatre, reconstructed
The launch choreography is worth recording, because it was choreography.
On 17 July, Moonshot released Kimi K3 — 2.8 trillion parameters, the model that briefly rattled US tech stocks and then suspended new subscriptions when demand outran capacity. On 18 July, an anonymous model called "kaleb" appeared on the Code Arena leaderboard, introducing itself as "Claude" — a distillation artifact — until the community identified it within a day by a quirk unique to Alibaba's Qwen tokenizer. On 19 July, during the World AI Conference in Shanghai, Alibaba confirmed it: kaleb was Qwen3.8-Max, previewed in stealth.
The Sunday announcement carried the "second only to Fable 5" claim and nothing to check it against. Alibaba's shares rose as much as 5.4 percent on the Monday anyway. The preview was purchasable through the Token Plan at ten percent of standard pricing — a real endpoint with real limits (a 983,616-token context window, 131,072 maximum output, thinking always on with three effort settings) and no paper trail.
That was the state of play until this morning. Announcing the claim on a Sunday and the evidence fifteen days later is a strategy, and it worked: two weeks of coverage ran on Alibaba's framing because there was nothing else to run on.
As an affiliate, we earn on qualifying purchases.
What the spec sheet actually says
The confirmed shape: 2.4 trillion total parameters, roughly 95 billion active per query, sparse mixture-of-experts, built on the Qwen3.5 architecture. Multimodal in — text, image, video — text out. The active-parameter figure is the number nobody had, and it settles the most important open question: this is a roughly-95B-compute model wearing a 2.4T coat, about 4 percent of the network firing per token.
The benchmark table, on Alibaba's own harness, splits cleanly into three stories.
Where it leads: Terminal-Bench 2.1 at 86.6 — ahead of both Claude Opus 4.8 and Claude Fable 5 at 84.6, behind only GPT-5.6 Sol at maximum effort at 88.8. PaperBench at 93.0, top of the table. IFBench at 82.8. The multimodal and agentic rows are where the model shines: OSWorld-Verified at 86.1, Parametric CAD Bench at 91.5, OmniDocBench 1.5 at 92.1. Alibaba also demonstrated the long-horizon pitch directly: the model reproduced all six main results of a research paper, then tested eighteen of its own ideas across four rounds and beat the paper's method on AIME24 by 2.7 points.
Where it trails, badly: SWE-bench Pro at 67.7 against Fable 5's 80.0. FrontierSWE at 73.5 against Fable 5's 88.8. These are not rounding errors — they are twelve- and fifteen-point gaps on exactly the deep software-engineering benchmarks that the "second only to" framing invites you to assume are covered.
Where it barely moved: GPQA Diamond at 92.6, up from Qwen3.7-Max's 92.4. The reasoning ceiling did not lift. What lifted — enormously — is agentic execution against its own predecessor: DeepSWE from 21.6 to 56.6, FrontierSWE from 40.7 to 73.5, JobBench from 31.3 to 53.4.
So the honest one-line read: "second only to Fable 5" is true on the rows Alibaba chose and false on the rows it didn't. This is not scandalous — every lab's launch table does this — but it means the claim that moved a share price is a claim about a subset. The clearest genuine achievement is different and less quotable: Alibaba took its own model from unusable to competitive on long-horizon agent work in one generation, mostly through RL-environment scaling, and published the trajectory.
As an affiliate, we earn on qualifying purchases.
The lane question
For this desk, the open-weight status is the operative fact, and it is now three-tiered.
The hosted API is live today, OpenAI- and DashScope-compatible — a base-URL change for anyone who wants to A/B it.
The 2.4T weights, due next week, are a gesture more than a deployment option. Whatever the licence turns out to be — and it is still unpublished, which matters, given that Qwen's open models have historically shipped Apache 2.0 while Kimi K3's licence carries revenue and attribution triggers — a 2.4-trillion-parameter checkpoint is a multi-node datacenter artifact. With the activated count at 95B, even aggressive quantisation puts the memory footprint far beyond any single machine. Nobody reading this self-hosts it. Its release is real openness and it is also a flag planted: the largest open-weight model ever shipped, if it ships.
Qwen3.8-27B is the row that belongs in the local-first lane. A 27B checkpoint distilled from whatever made the flagship's agentic numbers jump is precisely the shape of model that runs on a single high-memory machine as a daily driver — and precisely the class where this desk's inference actually happens. Whether the agentic gains survive the compression is the question that decides whether next week matters. There is no benchmark table for the 27B yet. There is a pattern, though: the interesting Qwen releases have consistently been the mid-size ones, and the flagship's job has been to make headlines that the 30B-class models then monetise in deployments.
As an affiliate, we earn on qualifying purchases.
Bull and bear
The bull case: the agentic-generation jump is real and internally consistent across a dozen rows; the RL-environment scaling story gives a mechanism, not just a number. The active-parameter disclosure and full table — even fifteen days late — is more than Kimi K3 shipped at launch. If the 2.4T weights land under Apache 2.0, the ceiling of what "open weight" means moves permanently, and every argument that frontier capability requires closed weights loses another data point. And the 27B sibling could be the best local agent model available, on hardware people already own.
The bear case: every number above is Alibaba's own harness, and agent benchmarks are the most harness-sensitive numbers in the field — independent testing of Kimi K3 already tempered that model's launch claims substantially, and there is no reason to expect this table survives contact with third parties intact. The deep-SWE gaps mean the flagship coding-agent use case — the one that pays — still belongs to Fable 5 by a wide margin. "Next week" is a promise from a company that spent two weeks not publishing a benchmark table it possessed. And the licence is still unwritten: until it exists, "going open-weight" is a press strategy, not a property of the model.
There is also a quieter structural point. Kimi K3, DeepSeek's 0731, and now Qwen3.8-Max arrived within seventeen days of each other, each timed against the others, each claiming adjacency to Fable 5 — a model the US temporarily placed under export controls, which has made "close to Fable 5" the season's universal marketing coordinate. When three labs in three weeks all measure themselves against the same restricted model, the benchmark being run is partly geopolitical. That contest is real, but it is not the same thing as your workload.
As an affiliate, we earn on qualifying purchases.
What I'll actually do
Nothing, until three things exist: the licence text, the 27B benchmark table, and one independent harness run of the flagship's agentic rows. The API is a base-URL change whenever a comparison is worth running; the 2.4T weights are not a deployment option for this fleet regardless of licence; and the 27B — the release that could genuinely change the daily-driver calculus here — is currently a name in a press release.
The preview phase of this launch ran on a slogan for fifteen days because nothing checkable existed. The GA phase should not get the same courtesy. The table is out; now the table gets tested.
Sources: Alibaba Qwen team release and benchmark table via MarkTechPost, 3 August 2026; The Decoder on the 95B active-parameter figure, weights timeline, and RL-scaling account, 3 August 2026; Bloomberg on the release and Fable 5 comparison, 3 August 2026; MarkTechPost and eesel coverage of the 19 July WAIC preview; Wan27 and community reporting on the "kaleb" stealth identification; contemporaneous coverage of Kimi K3's launch, licence terms, and independent evaluations. All Qwen3.8-Max performance figures are Alibaba's own harness; the open-weight licence was unpublished at time of writing. Point-in-time as of 3 August 2026. Not purchasing advice.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
