Alibaba Qwen3.8-Max: 2.4 Trillion Parameters and a Price Below Claude Opus 5

Category: Tech Deep Dives

Reviewed by the aifreetool Editorial Team — a group of full-time AI-tool researchers and writers who verify every product claim against primary sources and independent testing.
Last updated August 12, 2026.
We keep no affiliate relationship with the products covered here and earn nothing if you click through. Where a claim could not be verified, we say so.

On August 3, 2026, Alibaba dropped Qwen3.8-Max, the largest model it has ever shipped. Qwen3.8-Max is a 2.4-trillion-parameter sparse mixture-of-experts system that activates only 95 billion parameters per token, supports a 1-million-token context window, and ranks fifth on the Text Arena leaderboard while sitting second on Vision Arena. The headline is not just the size; it is the price. Alibaba is undercutting Claude Opus 5 on international API pricing, and it plans to release the weights next week. For a flagship-class model, that combination is rare.

What Qwen3.8-Max Actually Is

Alibaba Cloud Qwen3.8-Max Launch
Source: www.alibabacloud.com — https://www.alibabacloud.com/blog/alibaba-unveils-qwen3-8-max-its-largest-and-most-capable-flagship-model-to-date_603420

Qwen3.8-Max builds on the Qwen 3.5 architecture with a sparse MoE design paired with a hybrid attention mechanism. The 2.4 trillion total parameters make it one of the biggest publicly disclosed models, yet only 95 billion are activated for any given token. That keeps inference costs closer to a mid-sized dense model than to a raw 2.4-trillion-parameter monster.

The model is multimodal out of the box. It accepts text, images, and video, and it can output long-form reasoning with a context window of up to one million tokens. Alibaba is positioning it as an autonomous worker rather than a chatbot: it can ingest 200-page documents, watch 100-hour video streams, and operate inside desktop and browser environments through agent frameworks. The release was paired with QwenWork, an all-in-one workplace agent platform that bundles desktop, cloud, and collaboration agents into a single client. That combination suggests Alibaba wants to own the application layer, not just supply a model.

Pricing is the other lever. The flagship API is live on Alibaba Cloud Model Studio today, and Chinese media reports cited by DataNorth AI say international input pricing lands at roughly 40 percent of Claude Opus 5, with output pricing at roughly 24 percent. When a model that ranks fifth on Text Arena and second on Vision Arena is sold at that discount, closed labs have to respond.

If you are tracking the open-weights race, this belongs on your watchlist alongside Kimi K3 and Llama 5. You can browse comparable models on aifreetool.site/ai-engine-model.

Benchmarks: Where Qwen3.8-Max Leads and Trails

DataNorth AI Qwen3.8-Max Analysis
Source: datanorth.ai — https://datanorth.ai/news/alibaba-releases-qwen3-8-max

Alibaba published an extensive benchmark table, and independent reviewers at DataNorth AI have pulled out the signal from the noise. The model leads on research-oriented coding and multimodal agent tasks, but it still trails Claude Fable 5 on real-world software engineering.

BenchmarkQwen3.8-MaxClaude Fable 5GPT-5.6 Sol
PaperBench93.088.890.5
OSWorld-Verified86.185.083.2
SWE-bench Pro67.780.0n/a
FrontierSWE73.588.8n/a
Terminal-Bench 2.186.684.688.8

Two caveats matter. Every score is vendor-reported from Alibaba's own harness, so independent replication is still pending. And the gap on FrontierSWE and SWE-bench Pro shows that deep, production-grade engineering remains Claude Fable 5's turf. Still, the generation-over-generation jump inside Qwen's own family is real: DeepSWE 1.1 moved from 21.6 to 56.6, FrontierSWE from 40.7 to 73.5, and JobBench from 31.3 to 53.4.

The 16-Day Coding Demo That Got People's Attention

The most talked-about demonstration was not a benchmark chart. Alibaba gave Qwen3.8-Max an empty folder and a simple instruction: build a self-evolving agent framework. The model ran for roughly 16 days without human intervention, turning user requests into GitHub issues, assigning tasks to itself, writing code, running tests, and folding in community feedback. By the end of the run it had produced 265 commits, 127 pull requests, and 151 issues, and the resulting project, called oh-my-cli, was open-sourced on GitHub.

Other case studies were just as aggressive. The model reproduced a research paper end-to-end in about 125 hours of compute, ran 33 GPU training jobs, and then improved on the original method by 2.7 points on AIME24. In a chip-design sandbox it reduced a working circuit from 8,298 logic gates to 678 over roughly 500 iterations. In a one-year e-commerce simulation it turned a 100,000-yuan starting budget into 416,252 yuan, finishing 38 percent ahead of the runner-up.

Key Takeaways

  • Qwen3.8-Max is a 2.4T-parameter MoE model with 95B active parameters and a 1M-token context window.
  • It leads on research coding and multimodal agent benchmarks but still trails Claude Fable 5 on deep software engineering.
  • Alibaba plans to release the weights next week, making it the first Max-class Qwen model to go open-weight.
  • International API pricing is set well below Claude Opus 5, which pressures closed-model margins.

My Take / The Bottom Line

The Qwen3.8-Max launch is a pricing and distribution move as much as a technical one. Alibaba is betting that a near-frontier open-weight model, sold at a discount to Western closed APIs, will pull enterprise developers into its cloud. That strategy has worked before for Chinese hardware and telecom equipment, and it may work again for AI inference.

For most teams, the model is not yet a drop-in replacement for Claude Fable 5 on hard engineering tasks. The gap on SWE-bench Pro is real. But if you are building agents, evaluating multimodal pipelines, or simply want optionality outside the OpenAI-Anthropic duopoly, Qwen3.8-Max is now a credible alternative. My recommendation: test it on your own data when the weights land, benchmark the inference cost per task, and keep a second provider in production. The winners here are developers who refuse to be locked into a single frontier lab.

Frequently Asked Questions

How big is Qwen3.8-Max?

It has 2.4 trillion total parameters, but only 95 billion are activated per token through a sparse mixture-of-experts architecture.

When will the weights be released?

Alibaba says Qwen3.8-Max weights will be released next week, alongside a smaller Qwen3.8-27B checkpoint.

How does the pricing compare to Claude Opus 5?

Alibaba has priced the international API at roughly 40 percent of Opus 5 input pricing and 24 percent of output pricing, according to Chinese media reports cited by DataNorth AI.

What is QwenWork?

QwenWork is the workplace agent platform Alibaba launched alongside the model, integrating desktop agents, cloud agents, and enterprise collaboration tools.

Is Qwen3.8-Max better than Claude Fable 5?

It depends on the task. Qwen3.8-Max leads on research coding and multimodal agents; Claude Fable 5 still leads on production software engineering benchmarks like SWE-bench Pro and FrontierSWE.

FacebookXWhatsAppEmail