Kimi K3 is the 7th in our tests while it hit #1 on Arena Frontend Code at 1679 points!
We ran it the next day on our coding-agent repair harness against six peers: GPT-5.6 Sol, Fable 5, Grok 4.5, Opus 4.8, GLM-5.2, and Gemini 3.1 Pro.
It finished last of the 7 models with 53 of 67 attempts (79%).
So, How can both be true?
Read why the boards disagree (and the full scoreboard) here ↓↓
alphasignal.ai
On AlphaSignal’s signaldesk coding-agent run ( 2026-07-17-full-kimi-k3-launch , 13 tasks, 7 models, 488 attempts), Moonshot’s kimi-k3 resolved 53 of 67