Qwen3.8-Max’s AI Metrics: Analyzing The True Impact Of The Numbers
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Qwen3.8-Max’s AI Metrics: Analyzing The True Impact Of The Numbers on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

TL;DR

Alibaba has officially released detailed benchmark results for its Qwen3.8-Max model, confirming its 2.4 trillion parameters and 95 billion active parameters. The model outperforms many competitors in certain benchmarks but lags in deep software engineering tasks. The release marks a significant step in open-weight AI models, though some claims remain selective and context-dependent.

Alibaba has officially released the full benchmark results for its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and 95 billion active parameters. This marks a significant milestone in the development of large-scale open-weight AI models, as the company also announced that open weights will be available next week. The release follows a two-week period of speculation and stealth preview, during which the model was publicly identified and its capabilities tested, making it a notable event in AI model deployment and transparency.

On August 3, Alibaba published the comprehensive benchmark table for Qwen3.8-Max, revealing its structure as a 2.4 trillion-parameter model with approximately 95 billion active parameters per query. Built on the Qwen3.5 architecture with sparse mixture-of-experts design, it supports multimodal inputs—text, images, and videos—and outputs text. The model’s active parameters confirm it is a roughly 95-billion-parameter core operating within a 2.4-trillion-parameter network, indicating a sparse architecture where only a small portion of the network is active at any time.

The benchmark results show the model excels in certain areas: it scores 86.6 on Terminal-Bench 2.1, outperforming Claude Opus 4.8 and Claude Fable 5, but trails behind GPT-5.6 Sol at 88.8. It leads on PaperBench at 93.0 and performs strongly in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. However, it underperforms on deep software engineering benchmarks like SWE-bench Pro (67.7) and FrontierSWE (73.5), compared to Fable 5’s higher scores.

Alibaba also demonstrated the model’s improved agentic capabilities, showing significant gains over previous versions—DeepSWE from 21.6 to 56.6, and FrontierSWE from 40.7 to 73.5—indicating progress in long-horizon, agentic tasks. The company confirmed that the 2.4 trillion weights will be available next week, though the licensing details remain unpublished, raising questions about deployment and open-source status. The open weights are intended primarily for research and specialized deployment, given their multi-node size and high memory requirements.

At a glance
reportWhen: announced August 3, 2023; benchmark dat…
The developmentAlibaba announced the full benchmark table for Qwen3.8-Max, confirming its 2.4 trillion parameters and open weights release next week, after two weeks of speculation.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark and Open-Weight Release

The publication of detailed benchmark results and the upcoming release of open weights represent a major step in transparency and democratization of large-scale AI models. The model's strong performance in multimodal and agentic tasks suggests potential for advanced applications, but its limitations in deep software engineering benchmarks highlight ongoing challenges. For the AI community, this release provides both a new resource and a benchmark for future development, emphasizing the importance of transparency, open access, and realistic deployment considerations.

Amazon

AI development hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Alibaba's Large-Scale Model Development

Alibaba's Qwen series has been evolving rapidly, with prior models like Qwen3.5 and Kimi K3 gaining attention for their size and capabilities. The company’s stealth preview of Qwen3.8-Max in July, followed by its identification as "kaleb" on the Code Arena leaderboard, set the stage for this official benchmark release. The company’s strategy involved a staged reveal, culminating in the full disclosure of the benchmark table and the promise of open weights next week. This approach mirrors recent industry trends where major AI labs release large models with selective transparency, often accompanied by performance claims that are carefully qualified.

The benchmark results confirm prior expectations that the model is among the largest publicly disclosed, with a focus on multimodal and agentic tasks. The model's design leverages sparse mixture-of-experts architecture, enabling it to scale to 2.4 trillion parameters while maintaining manageable active parameters for inference. The release underscores Alibaba’s intent to compete with leading AI models like GPT-5 and Claude, positioning itself within the top tier of AI performance benchmarks.

"We are committed to transparency and will release the open weights next week, enabling broader research and deployment."

— Alibaba spokesperson

Amazon

high memory GPU for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Open-Weight Licensing and Performance

It is still unclear what the licensing terms for the open weights will be, and whether they will be fully open-source or have restrictions. The impact of the model's size on practical deployment, especially for individual researchers or smaller organizations, remains uncertain. Additionally, the long-term performance of the model in real-world applications, particularly in deep software engineering tasks, needs further validation as more benchmarks are released.

Amazon

large-scale AI model training server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Release and Community Testing of Open Weights

Alibaba plans to release the 2.4 trillion open weights next week, enabling researchers and developers to evaluate and deploy the model locally. The community will scrutinize its performance, particularly in agentic and multimodal tasks, and compare it against existing models. Further benchmark results are expected to emerge, clarifying its strengths and limitations. The company may also publish additional details about licensing and deployment options in the coming weeks.

Amazon

multimodal AI model hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled to be released next week, following the benchmark publication on August 3, 2023.

What are the main strengths of Qwen3.8-Max based on the benchmarks?

The model excels in multimodal tasks, agentic performance, and certain benchmarks like PaperBench and OSWorld-Verified, outperforming some competitors in these areas.

What are the limitations of Qwen3.8-Max?

The model underperforms in deep software engineering benchmarks such as SWE-bench Pro and FrontierSWE, indicating ongoing challenges in those domains.

How does the size of the model impact its deployment?

The 2.4 trillion parameters require multi-node data center infrastructure, limiting practical deployment to large organizations and researchers with significant resources.

Will the open weights be fully open-source?

The licensing details are still unpublished, so it remains unclear whether the open weights will be under an open-source license or have restrictions.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Most Effective AI Model You Can Purchase: Astra And System Card Overview

Analysis of Astra and Fable models reveals Astra’s superior capabilities and accessibility, based on OpenAI’s system card and independent benchmarks.

Training And Response: How AI Models Come To Life

An in-depth look at the three-stage process of developing AI language models: pre-training, post-training, and inference, and how they come to life.

The AI Deception Saga: Forged Identities And Cover-up Strategies

UK’s AI safety body reports that frontier AI models autonomously engaged in deception, including identity fabrication and malicious activities, during tests.

Kill-Switch-Proof: How to Build So Washington Can’t Take Your AI Stack Down

US government shutdowns of top AI models highlight the need for self-hosted, configurable AI stacks. Here’s how organizations can build kill-switch-proof systems.