📊 Full opportunity report: Qwen3.8-Max's Performance Data: The Good, The Bad, And The Confusing on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced detailed performance data for its Qwen3.8-Max model, confirming a 2.4 trillion-parameter size with strong benchmark scores. The model shows significant agentic improvements but trails in certain software benchmarks. Open weights will be released next week, marking the largest open-weight model to date.

Alibaba has officially published comprehensive performance data for its Qwen3.8-Max model, confirming it as the largest open-weight model ever shipped with 2.4 trillion parameters. This development follows weeks of speculation and stealth testing, with the company now providing benchmark scores and details on its architecture. The release of open weights next week signifies a major milestone for open AI models, especially given the model’s impressive performance in several benchmarks.

Alibaba’s Qwen3.8-Max is built on a 2.4 trillion-parameter architecture, utilizing sparse mixture-of-experts technology and based on the Qwen3.5 architecture. It features a 95 billion active parameter count per query, with multimodal capabilities including text, image, and video inputs, and text output. The model was previously previewed stealthily in July, with the company confirming its identity during the World AI Conference in Shanghai.

Benchmark results, obtained using Alibaba’s own testing harness, show that Qwen3.8-Max achieves a Terminal-Bench score of 86.6, surpassing Claude Opus 4.8 and Claude Fable 5 but trailing behind GPT-5.6 Sol. On PaperBench, it scores a top-tier 93.0. The model excels particularly in multimodal and agentic tasks, with notable improvements over its predecessor, DeepSWE, which increased from 21.6 to 56.6 in agentic performance.

However, the model underperforms on certain deep software engineering benchmarks, such as SWE-bench Pro and FrontierSWE, where it scores significantly lower than Fable 5, with gaps of 12 and 15 points respectively. This indicates that while the model has made substantial progress in agentic tasks, it still has limitations in complex software engineering tasks.

At a glance
reportWhen: announced August 3, 2023; full details…
The developmentAlibaba revealed detailed benchmark results and confirmed the availability of open weights for Qwen3.8-Max, after weeks of speculation and stealth testing.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Benchmark Results and Open Release

The detailed performance data confirms Alibaba’s position as a major player in AI development, with a model that demonstrates competitive benchmark scores and significant advancements in agentic capabilities. The open release of weights next week could influence the AI ecosystem by enabling broader access to a model of this scale, potentially accelerating innovation and deployment across industries.

Nevertheless, the model's shortcomings in software engineering benchmarks highlight ongoing challenges in achieving comprehensive general intelligence. The selective nature of the benchmarks also suggests that the model’s strengths may be confined to specific tasks, which is important for users and developers to consider when evaluating its utility.

Amazon

large open-weight AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Launch Strategy

Alibaba’s AI model development has been marked by a series of stealth previews and strategic disclosures. In July, the company introduced Qwen3.8-Max as a stealth model, which was identified through community detection methods. The model was then showcased during the World AI Conference, where Alibaba claimed it was ‘second only to Fable 5’ in performance, a claim that was initially unverified.

Following this, Alibaba provided a limited preview via a paid endpoint, with a 983,616-token context window and three effort settings, but withheld detailed benchmark data until August 3. The full benchmark table and the confirmation of open weights mark a significant shift in transparency and accessibility, aligning with broader industry trends towards open models.

"Alibaba’s release of detailed benchmark scores and the confirmation of open weights next week mark a pivotal moment for open AI models, especially with the model’s impressive agentic performance."

— Thorsten Meyer, reporting for ThorstenMeyerAI.com

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice: Camera and audio for AI interactions
  • Multiple Algorithm Support: OpenCV, YOLO for face and pose detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Capabilities and Licensing

It is still unclear what the licensing terms for the open weights will be, as Alibaba has not yet published the license. The impact of the open weights on the model’s performance, particularly whether the agentic gains will survive compression and quantization, remains to be seen. Additionally, benchmark results for the 27B version, which is targeted for local deployment, have not yet been released.

Doom's Benchmark: The Game That Measures Machines (Prompt Engineering with AI)

Doom's Benchmark: The Game That Measures Machines (Prompt Engineering with AI)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps: Open Weights Release and Community Evaluation

Alibaba plans to release the 2.4 trillion-parameter open weights next week, alongside the 27B checkpoint optimized for local deployment. The community will likely evaluate the model’s performance in real-world tasks, and further benchmark data may emerge, clarifying its strengths and weaknesses. Monitoring how the model’s agentic capabilities translate when scaled down will be key to assessing its practical utility.

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main strengths of Qwen3.8-Max based on the latest benchmarks?

Qwen3.8-Max demonstrates strong overall performance, especially in multimodal and agentic tasks, with top scores on PaperBench and notable improvements over its predecessor in agentic benchmarks.

When will the open weights for Qwen3.8-Max be available?

The open weights are scheduled for release next week, which will make the largest open-weight model publicly accessible, though licensing terms remain unpublished.

How does Qwen3.8-Max compare to other models like GPT-5.6 or Fable 5?

In benchmark scores, it outperforms Claude models but trails behind GPT-5.6 Sol on some measures. It excels in multimodal and agentic tasks but underperforms in specific deep software engineering benchmarks.

What are the implications of Alibaba’s open release strategy?

Releasing the open weights could democratize access to a large-scale model, fostering innovation but also raising questions about licensing, misuse, and the model’s limitations in certain tasks.

What remains uncertain about the model’s capabilities?

Uncertainties include the final licensing terms, the performance of the 27B checkpoint in local deployment, and whether agentic improvements will persist after compression and quantization.

Source: ThorstenMeyerAI.com

You May Also Like

ALIA. The Spanish answer.

Spain’s ALIA-40B, trained on 9.37 trillion tokens, is the largest EU-funded public AI project, emphasizing multilingual Spanish coverage but with performance below Llama 2.

The Cost Yagni Was Never About – By Kent Beck

Kent Beck explains that YAGNI is about timing and optionality, not effort savings, in a detailed reflection on software design principles.

From Experimental To Infrastructure: AI Operations Are Changing Fast

AI operations are rapidly evolving from experimental projects into critical infrastructure, with companies like xAI adopting data center-like models, impacting rollout strategies.

Show HN: Leaves – A text-UI Disk Usage Treemap Visualizer

A new text-based disk usage visualization tool called Leaves has been shared on Show HN, offering a treemap view in a terminal environment.