AI Rankings Shakeup? Qwen3.8-Max’s Latest Performance In Focus

📊 Full opportunity report: AI Rankings Shakeup? Qwen3.8-Max’s Latest Performance In Focus on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the full benchmark results for Qwen3.8-Max, revealing a 2.4 trillion-parameter model with 95 billion active parameters. The model outperforms several competitors on key benchmarks, marking a significant development in AI performance rankings. The open weights will be released next week, but questions remain about licensing and real-world deployment.

Alibaba has officially published the benchmark results for its Qwen3.8-Max model, confirming it as a 2.4 trillion-parameter, multimodal AI system with 95 billion active parameters. This marks the first time the model’s detailed performance data has been publicly available, challenging existing perceptions of AI model hierarchies and performance benchmarks. The release also confirms that open weights will be available next week, potentially influencing deployment strategies across the industry.

Alibaba’s Qwen3.8-Max, previously known only as a stealth preview and a slogan, has now been fully benchmarked on multiple standard tests. The model features a sparse mixture-of-experts architecture built on the Qwen3.5 foundation, with a total of 2.4 trillion parameters, but only approximately 95 billion active per query. This configuration allows it to deliver high performance while maintaining efficiency.

The benchmark results show that Qwen3.8-Max outperforms several leading models on key tests. It scores 86.6 on Terminal-Bench 2.1, surpassing Claude Fable 5 and Claude Opus 4.8, but trailing GPT-5.6 Sol at 88.8. On PaperBench, it achieved the top score of 93.0, and it demonstrated strong multimodal and agentic capabilities, with scores like 86.1 on OSWorld-Verified and 91.5 on Parametric CAD Bench. Notably, it improved its own agentic performance significantly compared to previous versions, indicating substantial advances in long-horizon reasoning and environment interaction.

Despite these gains, the model trails behind on some deep software-engineering benchmarks, such as SWE-bench Pro and FrontierSWE, with scores well below Fable 5. The model’s long-horizon reasoning capabilities, however, show marked improvement, with agentic scores jumping from early versions’ low teens to over 50 in some cases. The open weights are set to be released next week, but licensing details remain unpublished, raising questions about deployment and usage rights.

At a glance
updateWhen: announced August 3, 2023; benchmarks pu…
The developmentAlibaba officially published benchmark results for Qwen3.8-Max, confirming its 2.4 trillion parameters and competitive performance, shaking up AI rankings.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Impact of Qwen3.8-Max on AI Model Rankings

The publication of detailed benchmark results for Qwen3.8-Max signifies a major development in AI performance metrics, potentially reshaping industry rankings and competitive dynamics. Its high scores on key benchmarks demonstrate that Alibaba is making substantial progress in multimodal and agentic AI, challenging established leaders like OpenAI and Anthropic.

The open-weight release next week could accelerate adoption by researchers and developers, but licensing uncertainties and the model’s large size mean practical deployment may remain limited to data centers. The model's improved agentic capabilities suggest new possibilities for long-term reasoning and autonomous AI applications, which could influence future AI development and deployment strategies.

LLM Performance Evaluation: How to Build Automated Testing Pipelines, Benchmark Models, and Validate AI Applications Before Production

LLM Performance Evaluation: How to Build Automated Testing Pipelines, Benchmark Models, and Validate AI Applications Before Production

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Development of Qwen3.8-Max

Alibaba's Qwen series has evolved rapidly since its initial stealth preview in July 2023. The model was first hinted at during the World AI Conference in Shanghai, where Alibaba confirmed its existence and capabilities. Prior to the benchmark release, the model had been known only through a limited preview and a slogan claiming it was 'second only to Fable 5.' The benchmark results and full specifications now provide concrete data, ending weeks of speculation.

The development of Qwen3.8-Max builds on Alibaba's previous models, notably Qwen3.5, with significant architectural enhancements like sparse mixture-of-experts and multimodal integration. The company also introduced a smaller 27B variant, Qwen3.8-27B, optimized for local deployment and inference. The trajectory of Alibaba's AI development has focused on scaling, agentic reasoning, and multimodal capabilities, aligning with broader industry trends toward more capable and versatile models.

Prior to this release, models like Meta's Llama 2, OpenAI's GPT-4, and others had dominated performance benchmarks, but Alibaba's latest data suggests it is now a serious contender, especially in multimodal and long-horizon reasoning tasks.

"We are excited to share the detailed benchmark results and confirm the upcoming open release of our weights, which will facilitate broader research and deployment."

— Alibaba spokesperson

Kinupute Mini PC AI Server, AI Computing Workstation, AI MAX+ 395(126TOPS,16C/32T), Win-11 Pro, Radeon 8060S GPU, 128G LPDDR5X-8400, 8T M.2 SSD, 10G+2.5G LAN, Quad Screen, 4xM.2 PCIe 4.0 Slots, WiFi 7

Kinupute Mini PC AI Server, AI Computing Workstation, AI MAX+ 395(126TOPS,16C/32T), Win-11 Pro, Radeon 8060S GPU, 128G LPDDR5X-8400, 8T M.2 SSD, 10G+2.5G LAN, Quad Screen, 4xM.2 PCIe 4.0 Slots, WiFi 7

  • Purpose-Built for Local AI & LLM Inference: Optimized for on-premise AI and LLM tasks
  • 128GB LPDDR5X Memory: High-speed, massive onboard RAM for AI and video
  • Radeon 8060S GPU & XDNA 2 NPU: Powerful integrated graphics with AI acceleration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Licensing and Deployment

Details about the licensing terms for the open weights remain unpublished, raising questions about usage rights and commercial deployment. It is unclear whether the open weights will be fully open-source under permissive licenses like Apache 2.0 or subject to restrictions.

Additionally, the practical deployment of the 2.4 trillion-parameter model is limited by its size, requiring multi-node datacenter infrastructure. The performance of the smaller 27B variant in real-world tasks and its agentic capabilities remains to be tested outside benchmark environments.

Further, the long-term sustainability of the agentic improvements, especially under compression and real-world constraints, is still uncertain.

AI WORKSTATION GUIDE: A Practical Handbook for Developers, Data Scientists And Home AI Lab Builders on Hardware Selection, GPU Setup, LLM Deployment And Performance Optimization

AI WORKSTATION GUIDE: A Practical Handbook for Developers, Data Scientists And Home AI Lab Builders on Hardware Selection, GPU Setup, LLM Deployment And Performance Optimization

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s AI Model Strategy

Alibaba plans to release the open weights of Qwen3.8-Max next week, which will allow researchers and developers to evaluate its capabilities firsthand. The company is expected to publish licensing details simultaneously, clarifying usage rights.

In the coming months, attention will focus on how the smaller 27B variant performs in practical deployments, particularly in local or edge settings. The industry will also watch for further benchmarks and real-world applications to assess whether the agentic and multimodal improvements translate into tangible benefits.

Meanwhile, competitors will analyze Alibaba’s results to refine their own models, potentially leading to a new wave of AI innovations and performance benchmarks.

HOPLEX 20PCS Gundam Model Tool Kit Hobby Building Tools for Hobby Repairing

HOPLEX 20PCS Gundam Model Tool Kit Hobby Building Tools for Hobby Repairing

  • Versatile Hobby Use: Suitable for various crafts and models
  • Complete 20-Piece Set: Includes polishing, filing, cutting, and painting tools
  • Portable and Organized: All tools fit in a convenient carrying bag

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the key performance improvements of Qwen3.8-Max?

Qwen3.8-Max demonstrates superior scores on benchmarks like Terminal-Bench 2.1, PaperBench, and multimodal tasks, with significant gains in agentic reasoning and long-horizon tasks compared to previous versions.

When will the open weights for Qwen3.8-Max be released?

Alibaba announced that the open weights will be available next week, though the exact date and licensing terms are still to be confirmed.

How does Qwen3.8-Max compare to other leading models like GPT-4 or Fable 5?

On benchmarks like Terminal-Bench 2.1, Qwen3.8-Max surpasses some models but trails behind GPT-5.6 in certain areas. It excels in multimodal and agentic tasks, but its deep software engineering performance is still behind Fable 5.

What are the implications of Alibaba’s benchmark results for the AI industry?

The results suggest Alibaba is a serious competitor in high-end AI, especially in multimodal and agentic reasoning, which could influence future model development and industry rankings.

Source: ThorstenMeyerAI.com

You May Also Like

Apertus. The architectural template.

Apertus, developed by the Swiss AI Initiative, is a new open, multilingual, compliance-first AI model, offering a novel structural approach for European AI sovereignty.

Plant‑Based Bioart: Harnessing Photosynthesis for Artistic Expression

Discover how plant-based bioart leverages photosynthesis to create dynamic, living artworks that challenge traditional art boundaries and inspire ecological reflection.

How Documentation Builds Credibility in Experimental Art Practice

Building credibility in experimental art practice through documentation invites connection and reflection, but what transformative insights await those who embrace this journey?