Moonshot’s Kimi K3 tops front-end coding test, trails in hard math
Kimi K3 leads Code Arena’s front-end benchmark with a 1,679 Elo score, while Epoch AI puts it far behind OpenAI and Anthropic models on expert math.
By Renata Fuchs · Policy Reporter
· 3 min read
Moonshot’s Kimi K3 has taken the No. 1 position on Code Arena: Frontend, scoring 1,679 in a benchmark built around human preference ratings for front-end code. The same model looks much weaker on hard mathematics, with Epoch AI data showing about 39% accuracy on FrontierMath Tier 4, compared with scores near 90% for some OpenAI and Anthropic models.
The split result is a useful check on broad claims about frontier-model parity. Kimi K3 appears highly competitive, and in this case ahead, on a specific coding benchmark. It does not show the same performance on the most difficult expert-level math tasks tracked by Epoch AI.
Strong showing in front-end code
According to Code Arena: Frontend, Kimi K3’s 1,679 score puts it above Claude Fable 5, which is listed at 1,631, and GPT-5.6 Sol, which is listed at 1,618. The benchmark ranks models using human preference ratings, expressed through an Elo-style scoring system.
That puts Kimi K3 48 points ahead of Claude Fable 5 and 61 points ahead of GPT-5.6 Sol in the published ranking. Code Arena said the result marks the first time a Chinese model has led the benchmark.
For AI teams that use coding benchmarks as a proxy for developer-tool performance, the result is notable because front-end generation is a visible, commercially relevant workload. Model vendors and application builders are competing for use cases where code output can be judged quickly by users, especially in UI generation, prototyping and agent-assisted development.
The benchmark result still has limits. Code Arena: Frontend measures preference on front-end code outputs, not the full range of software engineering tasks. A leading score there does not by itself establish superiority in back-end reasoning, debugging across large codebases, systems design or other work developers assign to general-purpose coding assistants.
Math result shows a different ceiling
Epoch AI’s FrontierMath data points in the other direction. On Tier 4, described by Epoch AI as the hardest expert-level math category in the benchmark view cited, Kimi K3 reaches about 39% accuracy. Some OpenAI and Anthropic models are shown close to 90% on the same tier.
That gap matters because hard math benchmarks are often used as stress tests for abstract reasoning, long-horizon problem solving and reliability under expert-level constraints. Kimi K3’s coding result suggests strength in one practical domain, while the FrontierMath result indicates it remains well behind the strongest Western models on that benchmark.
The two data points also underline a recurring problem in model comparisons: a single leaderboard can overstate general capability. Kimi K3’s front-end score is a clear win on Code Arena’s terms. Epoch AI’s math data shows that the model’s performance is not uniform across demanding tasks.
Moonshot has drawn attention as Western AI labs face questions about whether compute scale still gives them a durable lead. These benchmark results do not settle that question. They show a Chinese model leading one human-preference coding test while lagging far behind OpenAI and Anthropic models on complex math.
This story draws on original reporting from The Decoder.