Qwen3.8 Max matches Claude Opus 4.8 as Kimi K3 posts a higher score
Artificial Analysis lists Qwen3.8 Max at 56, level with Claude Opus 4.8, but Kimi K3’s 57 comes with a lower reported task cost.
By Renata Fuchs · Policy Reporter
· 3 min read
Alibaba’s Qwen3.8 Max has reached a 56 on Artificial Analysis’ Intelligence Index, matching the reported score for Anthropic’s Claude Opus 4.8. In the Qwen3.8 Max vs Kimi K3 comparison, Kimi K3 remains one point ahead at 57 and carried a lower reported benchmark-task cost, though the underlying price figures vary by provider and measurement date.
Artificial Analysis lists Qwen3.8 Max as a proprietary model released in August, with a one-million-token context window and text, image and video inputs. Its 56 score places it ninth among 186 models in the displayed class, according to the benchmarking firm’s model page.
What does Qwen3.8 Max’s score mean for buyers?
The Intelligence Index is a composite rather than a direct measure of a single production workload. Version 4.1 combines nine evaluations, including GDPval-AA work tasks, agentic banking and terminal use, coding, scientific reasoning, knowledge tests, long-context reasoning and hallucination-related measures. A one-point gap between Qwen and Kimi should therefore not be read as a universal ranking for every deployment.
The score is a sharp improvement from Qwen3.7 Max’s reported 46. The Decoder, citing Artificial Analysis, also reported that Qwen3.8 Max achieved 1,739 Elo on GDPval-AA, ahead of Kimi K3’s 1,685, while Claude Opus 5 scored 1,852. That result shows how a model can lead on one component while trail on the composite index.
Why is Qwen’s reported task cost higher despite lower token rates?
Artificial Analysis lists Qwen’s API prices at $2 per million input tokens, $6 per million output tokens and $0.25 per million cached-input tokens. But The Decoder reported an estimated Intelligence Index cost of $1.14 per Qwen task, compared with $0.86 for Kimi K3. By that specific measure, Kimi’s figure is about 24.6% lower than Qwen’s, while Qwen is about 32.6% higher than Kimi.
That is a benchmark-cost comparison, not a blanket statement about all API spending. The Decoder reported that Qwen used 64 steps per GDPval-AA task versus 14 previously, and that input-token use rose 15-fold as the evaluation resent the full conversation history at each step. Artificial Analysis also characterizes Qwen as highly verbose: it produced 150 million output tokens over the index evaluation, versus a displayed median of 66 million.
The execution profile also has a speed trade-off. Artificial Analysis reports Qwen at 61.5 output tokens per second, below its cited same-class median of 71. Teams running long agent loops should test their own prompts, caching behavior and context handling rather than infer runtime or spend from published per-token rates.
There are data caveats. DeepInfra, a model provider, separately listed Kimi’s Artificial Analysis task cost at about $0.95, rather than $0.86, and used provider-specific API prices in its July comparison. Those differences make the figures snapshots shaped by provider, harness and model configuration.
Qwen’s gain also came with reported regressions against Qwen3.7 Max. The Decoder reported a two-point decline in AA-LCR, a long-context reasoning test, and a 10-point fall in AA-Omniscience. It said accuracy remained around 31%, while the reported hallucination rate rose from 23% to 40%. For operators, the practical conclusion is narrower than the headline: Qwen has closed the composite-score gap with Opus 4.8, but its reported steps and input-token use make workload-level evaluation necessary before treating lower token prices as lower total cost.
This story draws on original reporting from The Decoder.