Qwen3.8-Max’s AI Metrics: Analyzing The True Impact Of The Numbers

📊 Full opportunity report: Qwen3.8-Max’s AI Metrics: Analyzing The True Impact Of The Numbers on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has officially released detailed benchmark results for its Qwen3.8-Max model, confirming its 2.4 trillion parameters and 95 billion active parameters. The model outperforms many competitors in certain benchmarks but lags in deep software engineering tasks. The release marks a significant step in open-weight AI models, though some claims remain selective and context-dependent.

Alibaba has officially released the full benchmark results for its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and 95 billion active parameters. This marks a significant milestone in the development of large-scale open-weight AI models, as the company also announced that open weights will be available next week. The release follows a two-week period of speculation and stealth preview, during which the model was publicly identified and its capabilities tested, making it a notable event in AI model deployment and transparency.

On August 3, Alibaba published the comprehensive benchmark table for Qwen3.8-Max, revealing its structure as a 2.4 trillion-parameter model with approximately 95 billion active parameters per query. Built on the Qwen3.5 architecture with sparse mixture-of-experts design, it supports multimodal inputs—text, images, and videos—and outputs text. The model’s active parameters confirm it is a roughly 95-billion-parameter core operating within a 2.4-trillion-parameter network, indicating a sparse architecture where only a small portion of the network is active at any time.

The benchmark results show the model excels in certain areas: it scores 86.6 on Terminal-Bench 2.1, outperforming Claude Opus 4.8 and Claude Fable 5, but trails behind GPT-5.6 Sol at 88.8. It leads on PaperBench at 93.0 and performs strongly in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. However, it underperforms on deep software engineering benchmarks like SWE-bench Pro (67.7) and FrontierSWE (73.5), compared to Fable 5’s higher scores.

Alibaba also demonstrated the model’s improved agentic capabilities, showing significant gains over previous versions—DeepSWE from 21.6 to 56.6, and FrontierSWE from 40.7 to 73.5—indicating progress in long-horizon, agentic tasks. The company confirmed that the 2.4 trillion weights will be available next week, though the licensing details remain unpublished, raising questions about deployment and open-source status. The open weights are intended primarily for research and specialized deployment, given their multi-node size and high memory requirements.

At a glance
reportWhen: announced August 3, 2023; benchmark dat…
The developmentAlibaba announced the full benchmark table for Qwen3.8-Max, confirming its 2.4 trillion parameters and open weights release next week, after two weeks of speculation.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark and Open-Weight Release

The publication of detailed benchmark results and the upcoming release of open weights represent a major step in transparency and democratization of large-scale AI models. The model's strong performance in multimodal and agentic tasks suggests potential for advanced applications, but its limitations in deep software engineering benchmarks highlight ongoing challenges. For the AI community, this release provides both a new resource and a benchmark for future development, emphasizing the importance of transparency, open access, and realistic deployment considerations.

GPU Kernel Engineering for LLM Inference: CUDA, Triton, and Flash Attention Optimization for High-Throughput AI Production Systems (AI Infrastructure, Hardware & Compiler Engineering Series)

GPU Kernel Engineering for LLM Inference: CUDA, Triton, and Flash Attention Optimization for High-Throughput AI Production Systems (AI Infrastructure, Hardware & Compiler Engineering Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Alibaba's Large-Scale Model Development

Alibaba's Qwen series has been evolving rapidly, with prior models like Qwen3.5 and Kimi K3 gaining attention for their size and capabilities. The company’s stealth preview of Qwen3.8-Max in July, followed by its identification as "kaleb" on the Code Arena leaderboard, set the stage for this official benchmark release. The company’s strategy involved a staged reveal, culminating in the full disclosure of the benchmark table and the promise of open weights next week. This approach mirrors recent industry trends where major AI labs release large models with selective transparency, often accompanied by performance claims that are carefully qualified.

The benchmark results confirm prior expectations that the model is among the largest publicly disclosed, with a focus on multimodal and agentic tasks. The model's design leverages sparse mixture-of-experts architecture, enabling it to scale to 2.4 trillion parameters while maintaining manageable active parameters for inference. The release underscores Alibaba’s intent to compete with leading AI models like GPT-5 and Claude, positioning itself within the top tier of AI performance benchmarks.

"We are committed to transparency and will release the open weights next week, enabling broader research and deployment."

— Alibaba spokesperson

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950

  • System Compatibility: Measures 271 x 112 x 39 mm
  • Power Requirements: Requires 12V-2x6-pin connector
  • Customer Support: Contact us via Amazon for assistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Open-Weight Licensing and Performance

It is still unclear what the licensing terms for the open weights will be, and whether they will be fully open-source or have restrictions. The impact of the model's size on practical deployment, especially for individual researchers or smaller organizations, remains uncertain. Additionally, the long-term performance of the model in real-world applications, particularly in deep software engineering tasks, needs further validation as more benchmarks are released.

Amazon

large-scale AI model training server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Release and Community Testing of Open Weights

Alibaba plans to release the 2.4 trillion open weights next week, enabling researchers and developers to evaluate and deploy the model locally. The community will scrutinize its performance, particularly in agentic and multimodal tasks, and compare it against existing models. Further benchmark results are expected to emerge, clarifying its strengths and limitations. The company may also publish additional details about licensing and deployment options in the coming weeks.

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled to be released next week, following the benchmark publication on August 3, 2023.

What are the main strengths of Qwen3.8-Max based on the benchmarks?

The model excels in multimodal tasks, agentic performance, and certain benchmarks like PaperBench and OSWorld-Verified, outperforming some competitors in these areas.

What are the limitations of Qwen3.8-Max?

The model underperforms in deep software engineering benchmarks such as SWE-bench Pro and FrontierSWE, indicating ongoing challenges in those domains.

How does the size of the model impact its deployment?

The 2.4 trillion parameters require multi-node data center infrastructure, limiting practical deployment to large organizations and researchers with significant resources.

Will the open weights be fully open-source?

The licensing details are still unpublished, so it remains unclear whether the open weights will be under an open-source license or have restrictions.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

The City That Watches Itself: The Living Digital Twin, And The God’s-Eye View We’re Building

Cities are developing real-time digital replicas, combining sensors and AI to improve planning but raising surveillance concerns. Here’s what is confirmed and what remains uncertain.

Nvidia, CoreWeave, and Nebius: Inside the Circular Financing of the GPU Boom

An analysis of how Nvidia, CoreWeave, and Nebius are engaging in circular financing to fuel the GPU industry boom, with confirmed details and ongoing questions.

The Compounding Error Problem — Why 99.9% Alignment Decays to 60% in 500 Generations

Analysis of how 99.9% alignment accuracy declines rapidly over multiple AI generations, raising concerns about recursive self-improvement safety.

Acoustic Dampening, Placement, and the “Rig in the Closet” Setup

Learn effective strategies for reducing noise from AI workstations, including placement, acoustic dampening, and the ‘rig in the closet’ setup, with expert insights.