The Information Machine
Concluded·following since 4 Aug 2026·Day 6·16 sources·updated 13 Aug 2026

Alibaba's Qwen3-Max release

The gist

Qwen3.8-Max open weights scheduled for August 12; revenue-share terms still unfinalized

A 2.4-trillion-parameter open-weight model priced at $2/$6 per million tokens, with benchmark scores above several proprietary frontier models, puts competitive agentic capability within reach of self-hosted deployments. Alibaba's plan to charge large commercial users a revenue share on an open-weight release, if finalized, would be an unusual licensing condition in the open-source AI space.

The full picture

Alibaba's Qwen3.8-Max open weights are scheduled for public release on August 12, 2026, making it the first 'Max'-series Qwen model to receive open weights. A smaller companion dense model, Qwen3.8-27B, is also planned for the same release. The revenue-share rate Alibaba intends to charge large commercial users of the open-weight version has not been finalized, as negotiations are ongoing.

Qwen3.8-Max launched as a managed API on August 3, 2026. It is a sparse mixture-of-experts model with 2.4 trillion stored parameters that activates approximately 95 billion per token. The context window extends to 1 million tokens, with up to 131,072 output tokens per reply and a 262,000-token private thinking budget; it accepts text, image, and video input. API pricing is $2 per million input tokens and $6 per million output tokens, with cached reads at $0.17 per million.

On benchmarks, the model scores 86.1 on OSWorld-Verified, above GPT-5.6 Sol Max (83.2), Claude Fable 5 (85.0), and Gemini 3.1 Pro (76.2), and 93.0 on PaperBench, which DataCamp describes as the highest reported score on that benchmark. It scores 86.6 on Terminal Bench 2.1, between Claude Opus 4.8 (84.6) and GPT-5.6 Sol (88.8), and 92.6 on GPQA Diamond. Additional scores include $OneMillion-Bench (55.9), HealthBench (60.2), PLawBench (73.2), MRCR v2 256K (92.9), PRBench-Legal (57.6), and PRBench-Finance (58.3). The model outperformed Moonshot's Kimi K3 on several benchmarks.

Alongside the model, Alibaba introduced E-Commerce Bench, a 365-day simulation benchmark built on real, desensitized Taobao and Tmall transaction data covering 12 store types, 60 product categories, nearly 600 suppliers, and 7,000 products, with 152 fraudulent merchants embedded to test risk-control. In autonomous agent evaluations, the model reproduced a research paper's six findings by writing roughly 7,600 lines of code over five days, autonomously built a repository to 265 commits and 127 pull requests over sixteen days, optimized a chip design from 8,298 to 678 logic gates, and grew a simulated trading balance from 100,000 to 416,252 yuan.

Competitor Moonshot's Kimi K3 is larger at 2.8 trillion parameters, released open weights on July 27, and is priced at $3 per million input tokens and $15 per million output tokens. Alibaba stock rose 4.5% in New York premarket trading and 7% on the Hong Kong exchange following the Qwen3.8-Max announcement.

How it developed
13 August 2026

Alibaba confirmed August 12 as the open-weight release date for Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model launched as an API on August 3, and named a smaller companion, Qwen3.8-27B, for the same drop.

Alibaba's Qwen account also reported the model nearly doubled its score and climbed from rank 22 to rank 4 on Legal Research Bench in under three months. Alibaba stock rose 4.5% in New York and 7% in Hong Kong, while revenue-share negotiations for large commercial users of the open-weight version remain unresolved.

12 August 2026

Qwen3.8-Max Legal Research Bench ranking improved from 22nd to 4th, nearly doubling its score in under three months

9 August 2026

Alibaba released Qwen3-Max on August 8, a 2.4-trillion-parameter sparse mixture-of-experts model priced at $2.00 per million input tokens and available in Qwen Chat.

On Terminal Bench 2.1 it scored 86.6, between Claude Opus 4.8 (84.6) and GPT-5.6 Sol (88.8), and outperformed Moonshot's Kimi K3 on several benchmarks. Autonomous demonstrations included reproducing six research paper findings over five days and shrinking a chip design from 8,298 to 678 logic gates while cutting physical area by 81%.

First citedThe Neuron
8 August 2026

Detailed benchmark scores and autonomous agent evaluation results reported

Alibaba published Qwen3-Max on August 8, a 2.4-trillion-parameter sparse mixture-of-experts model activating roughly 95 billion parameters per token, with a 1 million token context window and pricing at $2.00 per million input tokens. It scored 86.6 on Terminal Bench 2.1, between Claude Opus 4.8 (84.6) and GPT-5.6 Sol (88.8), outperforming Moonshot's Kimi K3 on several benchmarks. In autonomous agent tests, it reproduced a research paper's six findings in code over five days and grew a simulated trading balance from 100,000 to 416,252 yuan.

7 August 2026

Reports emerge that Alibaba plans revenue-share requirement for large commercial users of the open-weight release

4 August 2026

Qwen3.8-Max becomes available in Qwen Chat with Unsloth support, priced at $2/M input and $6/M output tokens.

Sources
11 more sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free