
Moonshot shipped Kimi K3 on 16 July 2026: 2.8T parameters, 1M context, API live, weights promised by 27 July. AlphaSignal ran the model the next day on the same AlphaSignal coding-agent suite as six peers. Kimi K3 resolved 53 of 67 attempts (79%), finished last of seven models, cost $0.186 per successful fix, and averaged 702 seconds per attempt. Those numbers can all be true: they measure different jobs. This piece is about full-stack agent bug repair under held-out tests. What follows: how we score a fix, the full seven-model scoreboard, Arena #1 vs last-place repair, security and UI miss rates, cost and wall time per fix, and a Promote / Hold / Reject call. We use signaldesk v1, a ~7k-line Python and TypeScript app with planted bugs that don’t appear on public leaderboards. Each attempt starts in a fresh Docker sandbox with the network off. History is squashed, and held-out tests inject only at scoring time. A pass means visible tests, held-out tests, and the full regression suite all succeed. Agents also must not edit test files. Limits per attempt: 100 messages, 2M tokens, 30 minutes. Sampling stays at each provider's defaults. We rank models by resolve rate first, then by dollars per successful fix (total spend divided by solves). Cost accounting multiplies measured tokens by list rates. Kimi K3 uses Moonshot's public card: $3 per million cache-miss input, $15 per million output, $0.30 per million cache-hit input (platform pricing). Peer rates in the same run: GPT-5.6 Sol $5/$30, Fable 5 $10/$50, Opus 4.8 $5/$25, Grok 4.5 $2/$6, Gemini 3.1 Pro preview $2/$12, GLM-5.2 $0.42/$1.32. K3 ships with thinking always on and reasoning_effort max only (Kimi K3 quickstartAlphaSignal Signaldesk run: 13 tasks, 7 models, 488 attempts
Same launch window, other boards told a different story. Arena.ai put K3 at #1 on Frontend Code Arena with 1679 points, and Artificial Analysis scored it 57 on the Intelligence Index (#4 of 189).
How we scored a Kimi K3 agent fix
Run id: 2026-07-17-full-kimi-k3-launch. Thirteen tasks, seven models, 488 attempts in total, full report in sources below.
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves

