Qwen3.8-Max OSWorld benchmark claim puts Alibaba ahead of named rivals
Alibaba says Qwen3.8-Max scored 86.1 on OSWorld-Verified, but the computer-use result and long-running demos await independent testing.
By Wei-Lin Zhao · AI Correspondent
· 2 min read
Alibaba’s Qwen team has released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model, and says its Qwen3.8-Max OSWorld benchmark result exceeds the published scores of GPT-5.6 Sol Max and Fable 5. The model is available through QwenCloud, while Qwen says its weights will be released next week, with licensing terms still undisclosed.
Qwen reported an 86.1 score on OSWorld-Verified, compared with 83.2 for GPT-5.6 Sol Max, 85.0 for Fable 5 and 76.2 for Gemini 3.1 Pro. That places Qwen3.8-Max 2.9 points ahead of GPT-5.6 Sol Max and 1.1 points ahead of Fable 5 on the company’s comparison.
The announcement is a meaningful data point in the contest to build agents that can execute work rather than only respond to prompts. It is not, however, independent evidence that the result will generalize reliably to a production environment. The reported benchmark figures and Qwen’s demonstrations have not been broadly replicated by independent evaluators.
What does the Qwen3.8-Max OSWorld benchmark measure?
OSWorld-Verified evaluates computer-use agents interacting with desktop environments. For enterprise buyers, that category is relevant because it tests an agent’s ability to act through a graphical computing environment, rather than merely generate text or code in isolation.
Qwen3.8-Max has 2.4 trillion total parameters, 95 billion of them active, according to Qwen’s August 2 announcement. The company describes it as its most capable model so far and says it is built on the Qwen 3.5 architecture.
Qwen also published leading results for its model on several other evaluations, including 93.0 on PaperBench, 86.6 on TerminalBench 2.1, 69.0 on Vision2Web, 81.8 on LVBench and 77.8 on ERQA. Its reported profile is not a clean sweep: published reporting says an OpenAI model has the top reported score on SWE-Pro, while Opus 4.8 leads certain software-engineering tests and Agents’ Last Exam.
Qwen’s long-horizon coding claims remain first-party
Qwen says Qwen3.8-Max built the oh-my-cli project during an autonomous coding run lasting more than 10 days. As of July 30, the company said the repository had accumulated 265 commits, 127 pull requests and 151 issues after about 16 days of autonomous AI operation.
In a separate company-described research task, Qwen said the model worked for roughly five days to reproduce and improve a paper’s experiment, writing about 7,600 lines of code, taking more than 1,100 actions and running 33 GPU-training rounds. Those claims provide detail on the company’s test setup, not outside validation of model reliability.
QwenCloud lists API pricing of $2 per million input tokens and $6 per million output tokens. The company has called the coming release its first open-weight Qwen-Max model, but it has not stated the license. Until those terms arrive, enterprises cannot determine whether the weights are usable for their intended deployment and commercial conditions.
This story draws on original reporting from VentureBeat.