📊 Full opportunity report: AI Rankings Shakeup? Qwen3.8-Max’s Latest Performance In Focus on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba announced the full benchmark results for Qwen3.8-Max, revealing a 2.4 trillion-parameter model with 95 billion active parameters. The model outperforms several competitors on key benchmarks, marking a significant development in AI performance rankings. The open weights will be released next week, but questions remain about licensing and real-world deployment.
Alibaba has officially published the benchmark results for its Qwen3.8-Max model, confirming it as a 2.4 trillion-parameter, multimodal AI system with 95 billion active parameters. This marks the first time the model’s detailed performance data has been publicly available, challenging existing perceptions of AI model hierarchies and performance benchmarks. The release also confirms that open weights will be available next week, potentially influencing deployment strategies across the industry.
Alibaba’s Qwen3.8-Max, previously known only as a stealth preview and a slogan, has now been fully benchmarked on multiple standard tests. The model features a sparse mixture-of-experts architecture built on the Qwen3.5 foundation, with a total of 2.4 trillion parameters, but only approximately 95 billion active per query. This configuration allows it to deliver high performance while maintaining efficiency.
The benchmark results show that Qwen3.8-Max outperforms several leading models on key tests. It scores 86.6 on Terminal-Bench 2.1, surpassing Claude Fable 5 and Claude Opus 4.8, but trailing GPT-5.6 Sol at 88.8. On PaperBench, it achieved the top score of 93.0, and it demonstrated strong multimodal and agentic capabilities, with scores like 86.1 on OSWorld-Verified and 91.5 on Parametric CAD Bench. Notably, it improved its own agentic performance significantly compared to previous versions, indicating substantial advances in long-horizon reasoning and environment interaction.
Despite these gains, the model trails behind on some deep software-engineering benchmarks, such as SWE-bench Pro and FrontierSWE, with scores well below Fable 5. The model’s long-horizon reasoning capabilities, however, show marked improvement, with agentic scores jumping from early versions’ low teens to over 50 in some cases. The open weights are set to be released next week, but licensing details remain unpublished, raising questions about deployment and usage rights.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Impact of Qwen3.8-Max on AI Model Rankings
The publication of detailed benchmark results for Qwen3.8-Max signifies a major development in AI performance metrics, potentially reshaping industry rankings and competitive dynamics. Its high scores on key benchmarks demonstrate that Alibaba is making substantial progress in multimodal and agentic AI, challenging established leaders like OpenAI and Anthropic.
The open-weight release next week could accelerate adoption by researchers and developers, but licensing uncertainties and the model’s large size mean practical deployment may remain limited to data centers. The model's improved agentic capabilities suggest new possibilities for long-term reasoning and autonomous AI applications, which could influence future AI development and deployment strategies.

LLM Performance Evaluation: How to Build Automated Testing Pipelines, Benchmark Models, and Validate AI Applications Before Production
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background and Development of Qwen3.8-Max
Alibaba's Qwen series has evolved rapidly since its initial stealth preview in July 2023. The model was first hinted at during the World AI Conference in Shanghai, where Alibaba confirmed its existence and capabilities. Prior to the benchmark release, the model had been known only through a limited preview and a slogan claiming it was 'second only to Fable 5.' The benchmark results and full specifications now provide concrete data, ending weeks of speculation.
The development of Qwen3.8-Max builds on Alibaba's previous models, notably Qwen3.5, with significant architectural enhancements like sparse mixture-of-experts and multimodal integration. The company also introduced a smaller 27B variant, Qwen3.8-27B, optimized for local deployment and inference. The trajectory of Alibaba's AI development has focused on scaling, agentic reasoning, and multimodal capabilities, aligning with broader industry trends toward more capable and versatile models.
Prior to this release, models like Meta's Llama 2, OpenAI's GPT-4, and others had dominated performance benchmarks, but Alibaba's latest data suggests it is now a serious contender, especially in multimodal and long-horizon reasoning tasks.
"We are excited to share the detailed benchmark results and confirm the upcoming open release of our weights, which will facilitate broader research and deployment."
— Alibaba spokesperson

Kinupute Mini PC AI Server, AI Computing Workstation, AI MAX+ 395(126TOPS,16C/32T), Win-11 Pro, Radeon 8060S GPU, 128G LPDDR5X-8400, 8T M.2 SSD, 10G+2.5G LAN, Quad Screen, 4xM.2 PCIe 4.0 Slots, WiFi 7
- Purpose-Built for Local AI & LLM Inference: Optimized for on-premise AI and LLM tasks
- 128GB LPDDR5X Memory: High-speed, massive onboard RAM for AI and video
- Radeon 8060S GPU & XDNA 2 NPU: Powerful integrated graphics with AI acceleration
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Model Licensing and Deployment
Details about the licensing terms for the open weights remain unpublished, raising questions about usage rights and commercial deployment. It is unclear whether the open weights will be fully open-source under permissive licenses like Apache 2.0 or subject to restrictions.
Additionally, the practical deployment of the 2.4 trillion-parameter model is limited by its size, requiring multi-node datacenter infrastructure. The performance of the smaller 27B variant in real-world tasks and its agentic capabilities remains to be tested outside benchmark environments.
Further, the long-term sustainability of the agentic improvements, especially under compression and real-world constraints, is still uncertain.

AI WORKSTATION GUIDE: A Practical Handbook for Developers, Data Scientists And Home AI Lab Builders on Hardware Selection, GPU Setup, LLM Deployment And Performance Optimization
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Alibaba’s AI Model Strategy
Alibaba plans to release the open weights of Qwen3.8-Max next week, which will allow researchers and developers to evaluate its capabilities firsthand. The company is expected to publish licensing details simultaneously, clarifying usage rights.
In the coming months, attention will focus on how the smaller 27B variant performs in practical deployments, particularly in local or edge settings. The industry will also watch for further benchmarks and real-world applications to assess whether the agentic and multimodal improvements translate into tangible benefits.
Meanwhile, competitors will analyze Alibaba’s results to refine their own models, potentially leading to a new wave of AI innovations and performance benchmarks.

HOPLEX 20PCS Gundam Model Tool Kit Hobby Building Tools for Hobby Repairing
- Versatile Hobby Use: Suitable for various crafts and models
- Complete 20-Piece Set: Includes polishing, filing, cutting, and painting tools
- Portable and Organized: All tools fit in a convenient carrying bag
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What are the key performance improvements of Qwen3.8-Max?
Qwen3.8-Max demonstrates superior scores on benchmarks like Terminal-Bench 2.1, PaperBench, and multimodal tasks, with significant gains in agentic reasoning and long-horizon tasks compared to previous versions.
When will the open weights for Qwen3.8-Max be released?
Alibaba announced that the open weights will be available next week, though the exact date and licensing terms are still to be confirmed.
How does Qwen3.8-Max compare to other leading models like GPT-4 or Fable 5?
On benchmarks like Terminal-Bench 2.1, Qwen3.8-Max surpasses some models but trails behind GPT-5.6 in certain areas. It excels in multimodal and agentic tasks, but its deep software engineering performance is still behind Fable 5.
What are the implications of Alibaba’s benchmark results for the AI industry?
The results suggest Alibaba is a serious competitor in high-end AI, especially in multimodal and agentic reasoning, which could influence future model development and industry rankings.
Source: ThorstenMeyerAI.com