Artificial Analysis Rebuilds AI Leaderboards Around Real Legal and Medical Jobs

Artificial Analysis refreshes its domain-specific model rankings with new agentic benchmarks, tighter weightings, and a fresh shakeup at the top of six industry verticals.

·
·
Artificial Analysis Rebuilds AI Leaderboards Around Real Legal and Medical Jobs
Read4 min
TypeNews
  • Artificial Analysis released Capability Indices v1.1 covering six industry verticals with updated benchmark weights.
  • Agentic Tool Use added across Finance, Strategy, Legal, and Healthcare via AutomationBench-AA slices.
  • AA-Briefcase added to Agentic Knowledge Work in every index; Agentic Customer Interaction removed from four.
  • Engineering swaps GPQA Diamond for Terminal-Bench v4.0, shifting focus toward real shell execution.
  • Claude Fable 5.1 (max) leads all six indices; GPT-6 Astra (max) is second in four.
  • Open weights competitive: Kimi K3, DeepSeek V4.1 Flash, and GLM-5.3 land in top 10 across verticals.

Artificial Analysis has released version 1.1 of its Capability Indices, six domain-specific leaderboards designed to show how models perform on legal, financial, clinical, engineering, economic, and operational work. The update expands evaluations of multistep tool use and long-document reasoning, removes several older components, and recalculates the rankings.

Work tasks shape the scores

Artificial Analysis builds each index by mapping tasks from O*NET, the US occupational database, to relevant benchmarks. It then weights those benchmarks according to how frequently each capability appears in the corresponding jobs. Version 1.1 combines selected components from the broader Intelligence Index v4.3 with specialized evaluations across Finance and Accounting, Strategy and Ops, Legal, Healthcare and Medical, Engineering, and Economics.

Each vertical uses a different capability mix. Healthcare covers clinical knowledge, multistep knowledge work, reasoning across patient records, resistance to hallucination, clinical reasoning, and tool use. Engineering emphasizes technical knowledge, quantitative reasoning, task execution, and terminal use.

Tools and long documents gain weight

Version 1.1 expands agentic evaluations, which test whether a model can plan steps, call tools, and complete a workflow. The main changes are:

Change Affected indices Evaluation focus
Agentic Tool Use added Finance and Accounting, Strategy and Ops, Legal, Healthcare and Medical AutomationBench-AA slices covering finance, operations, and support workflows
AA-Briefcase added All six Multistep office tasks within Agentic Knowledge Work
GDP.pdf added Finance and Accounting, Strategy and Ops, Legal Reasoning over long documents
Agentic Customer Interaction removed Finance and Accounting, Strategy and Ops, Legal, Healthcare and Medical Customer-facing workflow performance no longer contributes to these composites
Engineering evaluations revised Engineering Terminal-Bench v4.0 replaces the previous terminal evaluation, while GPQA Diamond leaves the reasoning component
Long-Context Reasoning added Healthcare and Medical MLCR-AA tests medical reasoning across lengthy inputs

Frontier-model scores on GPQA Diamond have clustered near the benchmark’s ceiling, reducing its ability to separate leading systems. Terminal-Bench v4.0 instead measures whether a model can complete tasks in a command-line environment, including issuing commands, inspecting results, and recovering from errors.

One model leads every vertical

Claude Fable 5.1 under the site’s “max” configuration ranks first across all six indices. GPT-6 Astra (max) places second in Finance and Accounting, Strategy and Ops, Legal, and Engineering.

Models with downloadable weights also remain within the top 10. Their leading results by vertical are:

Model Leading verticals Overall rank
Kimi K3 (max) Finance and Accounting; Legal; Economics 8; 9; 7
DeepSeek V4.1 Flash (max) Strategy and Ops 7
GLM-5.3 (max) Healthcare and Medical; Engineering 6; 7

No open-weight model reaches the top five in a vertical. Rankings from sixth through ninth can still support a production shortlist when self-hosting, licensing, privacy, latency, hardware requirements, and deployment control influence the decision.

Execution now affects more scores

General intelligence leaderboards compress many abilities into one score. The Capability Indices narrow the comparison by weighting evaluations around occupational tasks, including document analysis, numerical reconciliation, citation accuracy, terminal work, and tool selection.

Adding AutomationBench-AA and AA-Briefcase increases the influence of multistep execution across the indices. GDP.pdf and MLCR-AA also give long-context performance a larger role in fields where models must process filings, contracts, or patient records. Artificial Analysis does not state a reason for removing Agentic Customer Interaction, so teams building customer-support systems should evaluate that capability separately.

These leaderboards remain benchmark aggregates, and their occupational weights may differ from a specific product’s workload. They also cannot capture every production constraint, including private data quality, tool reliability, prompt design, observability, and failure-recovery behavior.

Turn the ranking into a shortlist

The published scores already use version 1.1. Developers evaluating a domain-specific model can apply them through a focused selection process:

  1. Choose the closest vertical. Start with the index that best matches the product’s users and tasks.
  2. Inspect capability-level results. A composite rank can conceal weaknesses in tool use, long-context reasoning, or hallucination resistance.
  3. Match weights to the workload. Recalculate priorities when the application’s task mix differs from the occupational weighting.
  4. Compare operational constraints. Measure cost, latency, context limits, licensing, hosting options, and tool compatibility.
  5. Run application-specific tests. Use representative data, tools, prompts, and failure cases before choosing a production model.

A legal retrieval-augmented generation product that depends on analyzing long contracts should give AA-LCR v1.1 and GDP.pdf more weight than the headline rank. A finance agent that reconciles spreadsheets and calls external systems should focus on Agentic Tool Use and AA-Briefcase, then validate those results against its own workflow.

Comments

avatar