Aug 3, 2026
AI

Alibaba’s Qwen3.8-Max open weights are due next week

Alibaba has opened hosted access to Qwen3.8-Max and plans to release its weights next week, but its multi-day agent results remain unverified.

Wei-Lin Zhao

By Wei-Lin Zhao · AI Correspondent

· 3 min read

Alibaba’s Qwen3.8-Max open weights are due next week
Photo: The Decoder

Alibaba has made hosted access available for Qwen3.8-Max, its 2.4-trillion-parameter flagship model, and says Qwen3.8-Max open weights will be released next week. The launch matters because it would make this Alibaba’s first Qwen-Max-class model with downloadable weights, while the company argues it can sustain coding, research and multimodal workflows over hours or days rather than only answer a single prompt.

The Qwen team announced the model on August 2, with Alibaba Group publicizing it the following day. It is available through QwenCloud and an API now, according to Alibaba. That is separate from an open-weights release: developers can use the hosted product immediately, but the model files had not yet been released at the time of the announcement. The South China Morning Post described the move as Alibaba’s return to opening top-tier models after several recent flagship releases had remained proprietary.

What can Qwen3.8-Max do, and what has Alibaba proved?

Alibaba says Qwen3.8-Max is built on the Qwen3.5 architecture, has 2.4 trillion total parameters and activates 95 billion parameters for a given request. The difference reflects its sparse mixture-of-experts design: the full parameter count describes the model’s overall capacity, while only a portion is used for each query. Alibaba also claims a context window of up to 1 million tokens, and says the model can work with text, images and video, including documents longer than 200 pages and videos exceeding 100 hours.

The company’s central pitch is long-horizon agent work. Its most visible demonstration was an autonomous software project called oh-my-cli. Alibaba says Qwen3.8-Max worked on it for about 16 days without human help; by July 30, the project had produced 265 commits, 127 pull requests and 151 issues. The company says the system translated requests into issues, generated and tested code, then iterated on the result.

In a second company-run exercise, Alibaba gave the model a research paper but no starter code. It says the model spent about five days, or roughly 125 hours, writing about 7,600 lines of code and running 33 GPU training jobs. Alibaba says it reproduced six main findings from the paper, then tested 18 ideas across four rounds and improved on the paper’s method by 2.7 points on the AIME24 math benchmark.

Alibaba also reported a chip-design run of about 500 iterations, cutting a cryptographic circuit from 8,298 to 678 logic gates. In a separate fiscal-year e-commerce simulation based on anonymized Taobao and Tmall data, it said the model ended with 416,252 yuan from 100,000 yuan in starting capital, 38% ahead of the runner-up model. A simulation is not evidence of real-world retail performance.

These results, as well as Alibaba’s benchmark comparisons with other frontier models, are company-reported. The Decoder noted that independent verification remains pending. For developers, the release will test two distinct claims: whether the weights are useful enough to run and adapt outside Alibaba’s hosted services, and whether the model’s reported ability to complete extended tasks holds up in less controlled environments.

Alibaba’s technical announcement describes the intended release and its test setup. It does not establish independent performance validation, and Alibaba did not disclose in the cited materials how broadly developers will be able to reproduce the multi-day demonstrations.

This story draws on original reporting from The Decoder.

More from AI

All AI →