📊 Full opportunity report: Qwen3.8-Max’s AI Metrics: Analyzing The True Impact Of The Numbers on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba has officially released detailed benchmark results for its Qwen3.8-Max model, confirming its 2.4 trillion parameters and 95 billion active parameters. The model outperforms many competitors in certain benchmarks but lags in deep software engineering tasks. The release marks a significant step in open-weight AI models, though some claims remain selective and context-dependent.
Alibaba has officially released the full benchmark results for its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and 95 billion active parameters. This marks a significant milestone in the development of large-scale open-weight AI models, as the company also announced that open weights will be available next week. The release follows a two-week period of speculation and stealth preview, during which the model was publicly identified and its capabilities tested, making it a notable event in AI model deployment and transparency.
On August 3, Alibaba published the comprehensive benchmark table for Qwen3.8-Max, revealing its structure as a 2.4 trillion-parameter model with approximately 95 billion active parameters per query. Built on the Qwen3.5 architecture with sparse mixture-of-experts design, it supports multimodal inputs—text, images, and videos—and outputs text. The model’s active parameters confirm it is a roughly 95-billion-parameter core operating within a 2.4-trillion-parameter network, indicating a sparse architecture where only a small portion of the network is active at any time.
The benchmark results show the model excels in certain areas: it scores 86.6 on Terminal-Bench 2.1, outperforming Claude Opus 4.8 and Claude Fable 5, but trails behind GPT-5.6 Sol at 88.8. It leads on PaperBench at 93.0 and performs strongly in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. However, it underperforms on deep software engineering benchmarks like SWE-bench Pro (67.7) and FrontierSWE (73.5), compared to Fable 5’s higher scores.
Alibaba also demonstrated the model’s improved agentic capabilities, showing significant gains over previous versions—DeepSWE from 21.6 to 56.6, and FrontierSWE from 40.7 to 73.5—indicating progress in long-horizon, agentic tasks. The company confirmed that the 2.4 trillion weights will be available next week, though the licensing details remain unpublished, raising questions about deployment and open-source status. The open weights are intended primarily for research and specialized deployment, given their multi-node size and high memory requirements.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba's Benchmark and Open-Weight Release
The publication of detailed benchmark results and the upcoming release of open weights represent a major step in transparency and democratization of large-scale AI models. The model's strong performance in multimodal and agentic tasks suggests potential for advanced applications, but its limitations in deep software engineering benchmarks highlight ongoing challenges. For the AI community, this release provides both a new resource and a benchmark for future development, emphasizing the importance of transparency, open access, and realistic deployment considerations.

GPU Kernel Engineering for LLM Inference: CUDA, Triton, and Flash Attention Optimization for High-Throughput AI Production Systems (AI Infrastructure, Hardware & Compiler Engineering Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of Alibaba's Large-Scale Model Development
Alibaba's Qwen series has been evolving rapidly, with prior models like Qwen3.5 and Kimi K3 gaining attention for their size and capabilities. The company’s stealth preview of Qwen3.8-Max in July, followed by its identification as "kaleb" on the Code Arena leaderboard, set the stage for this official benchmark release. The company’s strategy involved a staged reveal, culminating in the full disclosure of the benchmark table and the promise of open weights next week. This approach mirrors recent industry trends where major AI labs release large models with selective transparency, often accompanied by performance claims that are carefully qualified.
The benchmark results confirm prior expectations that the model is among the largest publicly disclosed, with a focus on multimodal and agentic tasks. The model's design leverages sparse mixture-of-experts architecture, enabling it to scale to 2.4 trillion parameters while maintaining manageable active parameters for inference. The release underscores Alibaba’s intent to compete with leading AI models like GPT-5 and Claude, positioning itself within the top tier of AI performance benchmarks.
"We are committed to transparency and will release the open weights next week, enabling broader research and deployment."
— Alibaba spokesperson

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950
- System Compatibility: Measures 271 x 112 x 39 mm
- Power Requirements: Requires 12V-2x6-pin connector
- Customer Support: Contact us via Amazon for assistance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About Open-Weight Licensing and Performance
It is still unclear what the licensing terms for the open weights will be, and whether they will be fully open-source or have restrictions. The impact of the model's size on practical deployment, especially for individual researchers or smaller organizations, remains uncertain. Additionally, the long-term performance of the model in real-world applications, particularly in deep software engineering tasks, needs further validation as more benchmarks are released.
large-scale AI model training server
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Upcoming Release and Community Testing of Open Weights
Alibaba plans to release the 2.4 trillion open weights next week, enabling researchers and developers to evaluate and deploy the model locally. The community will scrutinize its performance, particularly in agentic and multimodal tasks, and compare it against existing models. Further benchmark results are expected to emerge, clarifying its strengths and limitations. The company may also publish additional details about licensing and deployment options in the coming weeks.

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
When will Alibaba release the open weights for Qwen3.8-Max?
The open weights are scheduled to be released next week, following the benchmark publication on August 3, 2023.
What are the main strengths of Qwen3.8-Max based on the benchmarks?
The model excels in multimodal tasks, agentic performance, and certain benchmarks like PaperBench and OSWorld-Verified, outperforming some competitors in these areas.
What are the limitations of Qwen3.8-Max?
The model underperforms in deep software engineering benchmarks such as SWE-bench Pro and FrontierSWE, indicating ongoing challenges in those domains.
How does the size of the model impact its deployment?
The 2.4 trillion parameters require multi-node data center infrastructure, limiting practical deployment to large organizations and researchers with significant resources.
Will the open weights be fully open-source?
The licensing details are still unpublished, so it remains unclear whether the open weights will be under an open-source license or have restrictions.
Source: ThorstenMeyerAI.com