
Moonshot AI just dropped Kimi K3, a 2.8-trillion-parameter open-weight model that is, by a significant margin, the largest open model ever released. On Artificial Analysis's new AA-Briefcase agentic benchmark, it scores second overall behind only Claude Fable 5 -- a remarkable result for an open-weight model. But that performance comes with a real catch: nearly an hour per task and a cost that rivals the most expensive proprietary models on the market.
What is AA-Briefcase, and why does it matter?
Most benchmarks test isolated capabilities -- math problems, coding puzzles, single-turn Q&A. AA-Briefcase is different. It evaluates models across four multi-week knowledge work projects, comprising thousands of input files and 91 tasks in total. Across the scenarios, models must complete realistic professional workflows in fields such as data science, product management, and corporate strategy. The outputs are actual deliverables -- spreadsheets, presentations, memos -- not multiple-choice answers.
Scoring combines three dimensions: binary rubric checks (did the model get the facts right?), pairwise analytical quality grading, and pairwise presentation quality grading. These are rolled into a single AA-Briefcase Elo number. It's the closest thing to a real-world office work test that exists at this scale.
Where Kimi K3 lands
At 2.8 trillion parameters it is the largest open-weight model released so far, and on independent testing it lands fourth among all frontier models -- trailing only Claude Fable 5 and GPT-5.6 Sol, and edging past Claude Opus 4.8. On AA-Briefcase specifically, the picture is even stronger. Kimi K3 achieves an Elo of 1543, a +727 improvement over its predecessor Kimi K2.6, and sits only behind Claude Fable 5 (1574). The full leaderboard context:
- Claude Fable 5: 1587 Elo (1st)
- Kimi K3: 1543 Elo (2nd)
- Claude Opus 4.8: 1356 Elo (3rd)
- GLM-5.2: 1266 Elo (4th)
- GPT-5.5 (xhigh): 1159 Elo (5th)
- Kimi K2.6: 809 Elo (19th)
That generational jump -- from 19th to 2nd -- is the headline. But the score breakdown reveals where K3 earns its rank and where it leaves points on the table.
Strong on analysis, weaker on polish
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves

