Make money doing the work you believe in

Kimi K3 is the 7th in our tests while it hit #1 on Arena Frontend Code at 1679 points!

We ran it the next day on our coding-agent repair harness against six peers: GPT-5.6 Sol, Fable 5, Grok 4.5, Opus 4.8, GLM-5.2, and Gemini 3.1 Pro.

It finished last of the 7 models with 53 of 67 attempts (79%).

So, How can both be true?

Read why the boards disagree (and the full scoreboard) here ↓↓

Jul 17
at
5:39 PM
Relevant people

Log in or sign up

Join the most interesting and insightful discussions.