Aug 13, 2026
AI

Ling 3.0 Flash tops Artificial Analysis’ medium open-weights benchmark

Ling 3.0 Flash ranks first among 63 medium open-weights models, though its high token use complicates the cost claim.

Renata Fuchs

By Renata Fuchs · Policy Reporter

· 3 min read

Ling 3.0 Flash tops Artificial Analysis’ medium open-weights benchmark
Photo: The Decoder

Ling 3.0 Flash benchmark results place the model first among 63 medium-sized open-weights models tracked by Artificial Analysis. The model scored 38 on the provider’s Intelligence Index, while its reported 124 billion total parameters and 5.1 billion active parameters put it at the upper end of the category. The result is a qualified leaderboard win, not a general finding that the model is the best choice for every workload.

Artificial Analysis, which lists Ling 3.0 Flash as an August 2026 release from InclusionAI, defines its medium open-weights class as models with 40 billion to 150 billion total parameters. The model is a text-in, text-out reasoning model with a 262,000-token context window. Artificial Analysis says a non-reasoning variant may also exist.

What does the Ling 3.0 Flash benchmark ranking mean?

The 38-point result is based on Artificial Analysis Intelligence Index v4.1.1, a composite of nine tests: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. Those tests cover areas including agentic work, tool use, coding, reasoning, knowledge and long-context performance.

A composite score makes the ranking useful for a broad comparison, but it does not prove performance on a company’s prompts, tools or quality threshold. Teams considering the model should test representative tasks and failure cases rather than treating one leaderboard position as a deployment decision. That is the distinction between a benchmark result and evaluating AI models for the work they will actually do.

The 124 billion figure also needs context. Total parameters describe the full model, while active parameters are those used per token during inference. Artificial Analysis lists 5.1 billion active parameters per token for Ling 3.0 Flash, a figure relevant to its operating profile but not a standalone measure of quality or infrastructure cost.

How fast and expensive is Ling 3.0 Flash?

Artificial Analysis measured output speed at 393.6 tokens per second. It lists API pricing of $0.07 per million input tokens and $0.22 per million output tokens, with an 80% cache discount, and estimates $0.04 as the weighted cost per Intelligence Index task. The firm says it spent $72.81 to run the model through the full Intelligence Index evaluation.

Those prices come with a material token-use caveat. Ling 3.0 Flash produced 240 million output tokens in the index evaluation, compared with a 57 million median among comparable models, according to Artificial Analysis. The provider ranks it 14th of 63 on verbosity and characterizes it as very verbose.

For operators, that means low published token rates do not automatically produce the lowest bill or shortest completion time. A model that uses substantially more output tokens can offset an input or output price advantage on complex tasks. Artificial Analysis nevertheless describes Ling 3.0 Flash as notably fast and reasonably priced against similarly sized open-weights models.

The available evidence supports the model’s category-specific rank and measured operating metrics. It does not establish multimodal support, general coding superiority, local deployment requirements, licensing terms or performance on any particular production workload.

This story draws on original reporting from The Decoder.

More from AI

All AI →