Arena #1 Kimi K3 Scored 79% on our Repair Bench

AlphaSignal Signaldesk run: 13 tasks, 7 models, 488 attempts

·
·
Arena #1 Kimi K3 Scored 79% on our Repair Bench
AuthorAdham Khaled
Read2 min
  • On AlphaSignal’s signaldesk coding-agent run (2026-07-17-full-kimi-k3-launch, 13 tasks, 7 models, 488 attempts), Moonshot’s kimi-k3 resolved 53 of 67 attempts (79%), last place, at $0.186 per successful fix and 702s average wall time.
  • Same launch window, Arena Frontend Code ranked K3 #1 at 1679 points (up 17 places from Kimi K2.6 at #18), while Artificial Analysis scored it 57 on the Intelligence Index (#4 of 189) with GDPval-AA v2 Elo 1668.
  • Against peers on the same harness, resolve gaps were large: GPT-5.6 Sol 100% (70/70) at $0.183 and 64s, Grok 4.5 99% at $0.074 and 46s, Fable 5 99% at $0.411, GLM-5.2 94% at $0.019, with only Gemini (80%) near K3’s bottom band.
  • Category resolve on that run hit logic 100% and performance 90%, but security only 69% (11/16) and web UI 70% (7/10), including 3/5 on import injection and admin auth bypass, while K3 burned ~115k tokens per attempt with 56% reasoning share and 3 over-edits (highest in the field).
  • List pricing is $0.30 cache-hit / $3 cache-miss / $15 output per 1M tokens with reasoning_effort max only, so the article’s call is Hold for signaldesk-style repair, Reject for unattended security patches, and Promote-to-trial for frontend work, not a single global ranking.

Moonshot shipped Kimi K3 on 16 July 2026: 2.8T parameters, 1M context, API live, weights promised by 27 July.

AlphaSignal ran the model the next day on the same AlphaSignal coding-agent suite as six peers.

Kimi K3 resolved 53 of 67 attempts (79%), finished last of seven models, cost $0.186 per successful fix, and averaged 702 seconds per attempt.

Same launch window, other boards told a different story. Arena.ai put K3 at #1 on Frontend Code Arena with 1679 points, and Artificial Analysis scored it 57 on the Intelligence Index (#4 of 189).

Those numbers can all be true: they measure different jobs. This piece is about full-stack agent bug repair under held-out tests.

What follows: how we score a fix, the full seven-model scoreboard, Arena #1 vs last-place repair, security and UI miss rates, cost and wall time per fix, and a Promote / Hold / Reject call.

How we scored a Kimi K3 agent fix

We use signaldesk v1, a ~7k-line Python and TypeScript app with planted bugs that don’t appear on public leaderboards.

Run id: 2026-07-17-full-kimi-k3-launch. Thirteen tasks, seven models, 488 attempts in total, full report in sources below.

Each attempt starts in a fresh Docker sandbox with the network off. History is squashed, and held-out tests inject only at scoring time.

A pass means visible tests, held-out tests, and the full regression suite all succeed. Agents also must not edit test files.

Limits per attempt: 100 messages, 2M tokens, 30 minutes. Sampling stays at each provider's defaults.

We rank models by resolve rate first, then by dollars per successful fix (total spend divided by solves).

Cost accounting multiplies measured tokens by list rates. Kimi K3 uses Moonshot's public card: $3 per million cache-miss input, $15 per million output, $0.30 per million cache-hit input (platform pricing).

Peer rates in the same run: GPT-5.6 Sol $5/$30, Fable 5 $10/$50, Opus 4.8 $5/$25, Grok 4.5 $2/$6, Gemini 3.1 Pro preview $2/$12, GLM-5.2 $0.42/$1.32.

K3 ships with thinking always on and reasoning_effort max only (Kimi K3 quickstart). On this run Fable 5 and Opus 4.8 show 0% reasoning share of output, matching Anthropic defaults with extended thinking off.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves