Artificial Analysis' Optima Lets Any Team Build Custom AI Benchmarks

Artificial Analysis launches Optima, letting teams build custom model benchmarks from their own data, agent traces, or plain-text descriptions — with cost and speed tracked alongside quality.

·
·
Read2 min
SubtopicCode Agents
  • Optima launched: Artificial Analysis releases a self-serve platform for building custom model benchmarks from your own data or agent traces.
  • Three input modes: Upload a dataset, import agent traces from Arize/Braintrust/Langfuse, or just describe your use case in plain text.
  • Research-grade grading: Uses the same pairwise judging methodology as AA-Briefcase and GDPval-AA; rubric grading at $0.25/criterion, pairwise at $0.75/match.
  • Cost + speed tracked: Every benchmark run surfaces Cost per Task and Time per Task alongside quality scores, enabling 10x cost-reduction comparisons.
  • Bring your own agent: Custom agent stacks can compete against frontier models in the same benchmark run over HTTP.
  • Available now: Usage-based pricing with no markup on raw model token costs; no subscription required.

Artificial Analysis, the independent benchmarking firm best known for its continuously updated LLM leaderboards and evaluations like AA-Briefcase and GDPval-AA, has launched Optima , a platform that lets anyone build a custom benchmark tailored to their own tasks, data, and use cases. The pitch is simple: standardized public benchmarks tell you which model is generally capable, but they cannot tell you which model is right for your finance agent, legal assistant, or image classification pipeline.

The timing is not accidental. LLM benchmarks in 2026 are necessary but insufficient. MMLU has saturated above 90%. HumanEval suffers from training data contamination. SWE-Bench scores vary 25 percentage points depending on scaffolding. No single benchmark predicts production performance reliably. Optima is Artificial Analysis' answer to that gap.

Three ways in, one leaderboard out

Optima offers three distinct entry points for building a benchmark, designed to meet teams wherever their data already lives:

  • Upload a dataset , bring an existing evaluation set from your own files or from Hugging Face directly.
  • Import agent traces , pull recorded agent sessions from platforms like Arize, Braintrust, or Langfuse. The LLM observability market is estimated at $2.69B in 2026 , and these are the tools where production traces already live for most teams.
  • Describe your use case , give Optima a plain-text description plus a few example inputs and outputs, and a build agent drafts the tasks and rubrics for you.

There is also an IDE integration: install the Optima skill and it can pull context directly from your coding environment and previous sessions to scaffold a benchmark without leaving your editor.

The grading engine is the real differentiator

Building the task set is only half the problem. Grading at scale is where most custom eval efforts fall apart. Optima brings two grading modes that mirror Artificial Analysis' own research-grade evaluations:

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves