Together AI Says Open-Source Kimi K3 Beats Claude Fable 5.1 on Legal Tasks

Together AI highlights that Moonshot's open-weight Kimi K3 outscores Anthropic's newer Claude Fable 5.1 by 60% on the hard slice of Harvey's autonomous legal agent benchmark.

·
·
  • Together AI reports Kimi K3 scores 60% higher than Claude Fable 5.1 on Harvey LAB-AA's hard autonomous legal tasks.
  • Kimi K3 already led the full LAB-AA leaderboard at 26.7% all-pass vs Fable 5's 14.2%.
  • LAB-AA tests 1,200+ real law-firm tasks across 24 practice areas with strict all-pass rubrics.
  • Fable 5.1 just launched in Harvey with major litigation and document-analysis gains over Fable 5.
  • K3 is a 2.8T-parameter open-weight MoE with 1M context, downloadable from Hugging Face.
  • Harvey's Tenet already post-trained K3 to nearly double LAB hold-out task completion.

Together says Kimi K3 leads Claude Fable 5.1 on hard legal-agent tasks

Inference provider Together AI says Moonshot AI’s open-weight Kimi K3 scored roughly 60% higher than Anthropic’s Claude Fable 5.1 on the hard subset of Harvey LAB-AA, an autonomous legal-work benchmark. The result matters for teams choosing legal-agent models because Harvey recently deployed Fable 5.1 for customers, while K3 can be downloaded, self-hosted, and post-trained under Moonshot’s license.

The reported 60% figure describes a relative lift, not a 60-percentage-point gap. The figures presented here omit the two absolute hard-split scores, task count, confidence intervals, and complete evaluation settings, which limits independent verification and prevents direct comparison with the overall LAB-AA leaderboard.

Legal AI company Harvey developed the Legal Agent Benchmark around more than 1,200 law-firm tasks spanning 24 practice areas. Tasks include reviewing matter files and producing memos, contract redlines, and presentations. Each submission is graded against an expert rubric, and the all-pass metric awards success only when every criterion passes.

Artificial Analysis’s LAB-AA implementation places an agent in a sandbox where it must inspect matter documents, use tools, and produce a finished deliverable. Together’s hard split concentrates on longer, multi-step matters that require planning, file handling, drafting, and revision without human intervention.

One headline, two score sets

Public overall LAB-AA results already place K3 ahead of several proprietary models:

Overall Artificial Analysis Harvey LAB-AA scores cited by Together
Model All-pass rate
Kimi K3 26.7%
Claude Fable 5 14.2%
Grok 4.5 13.3%

Claude Fable 5.1 is absent from that overall table. Together’s newer 60% claim covers only the hard subset, so developers should keep the two result sets separate when comparing models.

Harvey’s early Fable 5.1 testing found its largest gains over Fable 5 in litigation and dispute resolution, with strong results in data privacy, cybersecurity, and capital markets. Harvey also reported that the model retained quality at its lowest reasoning setting while using roughly half as many output tokens. Those internal findings use a different evaluation scope from Together’s hard-split comparison.

Cost strengthens Together’s case

Together sells K3 inference, giving it a commercial interest in comparisons that favor the model. Its earlier DeepSWE analysis compared K3 with Claude Fable 5 on software-engineering tasks. Pass@k measures whether at least one of k attempts solves a task.

DeepSWE results reported by Together
Metric Kimi K3 Claude Fable 5
Pass@1 68.5% 69.9%
Pass@2 82.0% 80.2%
Pass@4 89.4% 88.5%
Cost per rollout $4.65 $13.41

Fable 5 led on the first attempt, while K3 led when the harness allowed two or four attempts. Together calculated that K3 completed 2.8 times as many tasks per dollar. The LAB-AA claim extends its economic argument from coding agents to document-heavy legal workflows.

Open weights at 1.56 TB

Moonshot describes K3 as a 2.8-trillion-parameter model with vision support and a one-million-token context window. Its Kimi Delta Attention and Attention Residuals architectures are designed to support long-context processing. The model weights are available on Hugging Face.

The release is a 1.56 TB download, before accounting for runtime memory, key-value caches, redundancy, and serving overhead. Self-hosting therefore requires substantial distributed infrastructure or aggressive quantization. The custom Kimi K3 License governs use and redistribution, and teams should review its terms before deployment.

  • Keep a human approval gate. K3’s leading 26.7% all-pass score means 73.3% of evaluated tasks failed at least one rubric criterion.
  • Run task-specific evaluations. K3 ranks first on Artificial Analysis’s Harvey LAB-AA implementation, fourth on Vals AI’s implementation, and third on Vals AI’s legal-research benchmark. Harnesses and task distributions can change the ranking.
  • Freeze the evaluation stack. Record the model revision, tokenizer, quantization, system prompt, tools, reasoning level, token limits, attempt count, document set, and scoring rubric.
  • Treat the harness as part of the system. Harvey’s harness raised K3’s APEX score from 58.8% to 67.5%, an 8.7-percentage-point gain.
  • Validate post-training on held-out work. Harvey reports that Tenet, its post-trained model, completed nearly twice as many held-out LAB tasks and 20% more LAB Contracts tasks than base K3.
  • Budget for operational controls. Production legal systems need privilege boundaries, access controls, source citations, retention policies, audit logs, and review workflows alongside model inference.

What would settle the comparison

Together’s claim supports including K3 in a controlled evaluation against Fable 5.1. Production selection requires matched prompts, tools, documents, token budgets, attempt counts, and review rules, followed by analysis of latency, total serving cost, licensing, and security. Publishing the absolute hard-split scores, sample size, confidence intervals, and harness configuration would make the reported 60% lead independently testable.

Trending
  • No trending articles

Comments

avatar

Next Reads