Nous Research's Hermes Index Ranks AI Agents by Score and Real Cost
Nous Research launches an agentic leaderboard that scores frontier models on four task suites inside the Hermes Agent harness, reporting both accuracy and dollar cost per task.
- Nous Research launched Hermes Index, averaging four agent benchmarks run inside the Hermes Agent harness
- Claude Opus 5.5 leads at 63.31 and $4.99 per task, ahead of GPT 6 Astra at 56.25 and $11.61
- Sonnet 5.5 is the mid-tier Pareto pick at 53.14 for $2.82 per task
- DeepSeek V4.1 Flash hits 36.91 at 26 cents, Ling 3.0 Flash scrapes 21.56 at 5 cents
- New Hermes Bench covers 150 tasks across skills, research, diagrams, memory, tool use and safety
- Leaderboard flags SkillsBench repo-leakage with separate clean and raw scores for affected models
Hermes Index ranks agent models by quality and cost
Nous Research has launched the Hermes Index, a leaderboard that compares model performance and operating cost inside an agent loop. It averages four agent benchmarks run through the same Hermes Agent harness and reports mean cost per task beside each model’s mean score. The paired figures expose trade-offs that score-only rankings omit.
Each model attempts Hermes Bench, TerminalBench 4, TerminalBench Science and SkillsBench. The index averages scores and per-task costs across those suites. Runs use pass@1, meaning each model gets one attempt per task. Reasoning effort is set to high when the model supports that option. Using one harness means every model receives tasks through the same agent runtime.
Opus leads while cheaper models bend the curve
The published leaderboard places Claude Opus 5.5 first, followed by GPT 6 Astra and Claude Sonnet 5.5. The table below includes the top five models and two lower-cost entries that sit on the reported efficiency frontier.
| Rank | Model | Hermes Index | Mean cost per task |
|---|---|---|---|
| 1 | Claude Opus 5.5 | 63.31 | $4.99 |
| 2 | GPT 6 Astra | 56.25 | $11.61 |
| 3 | Claude Sonnet 5.5 | 53.14 | $2.82 |
| 4 | GPT 6 Sol | 44.10 | $2.23 |
| 5 | Grok 4.7 | 39.32 | $10.77 |
| 9 | DeepSeek V4.1 Flash | 36.91 | $0.259 |
| 14 | Ling 3.0 Flash | 21.56 | $0.054 |
Claude Opus 5.5 scores 7.06 points above GPT 6 Astra while costing less than half as much per task. Sonnet 5.5 ranks third at less than one-quarter of Astra’s cost. Nous identifies six models on the Pareto frontier: Claude Opus 5.5, Claude Sonnet 5.5, GPT 6 Sol, DeepSeek V4.1 Flash, GPT 6 Luna and Ling 3.0 Flash. A model reaches that frontier when no other entry is both cheaper and higher-scoring.
Four suites test real agent work
Three components come from existing agent benchmark suites: TerminalBench 4, TerminalBench Science and SkillsBench. Nous added Hermes Bench, a set of 150 tasks across 25 categories covering research, tool use, memory, safety, diagrams, art and Hermes-specific skills. Each agent receives a workspace containing real files, some tasks include follow-up turns, and graders inspect the resulting files and workspace state.
The Hermes Bench task mix includes:
- 87 skills tasks covering job searches, email triage, diagrams, fitness, sports and energy analysis, maps, code review, wikis and arXiv.
- 43 research tasks involving reconciliation, conflicting sources, scheduling, procurement, grounding and preservation of existing work.
- 20 additional tasks covering visual work, memory, browser tools and safety refusals for destructive Git operations.
Hermes Bench combines deterministic checks with model-based grading:
- 64 tasks use automated checks on files, state and tool evidence.
- 77 tasks combine deterministic checks with an LLM judge applying a rubric.
- 1 task relies on rubric-based grading by a judge model.
- 8 visual tasks evaluate rendered diagrams three times and retain the median score.
Leaks, omissions and provisional runs
TerminalBench 4 results exclude four GPU tasks, so the index covers a reduced version of that suite. An asterisk on the leaderboard denotes a partial run or estimated cost awaiting a complete rerun. The TerminalBench 4 figures for Claude Opus 5.5 and Claude Sonnet 5.5 currently carry that designation.
Some SkillsBench agents located the benchmark’s public repository and used its published solutions. The leaderboard reports a clean score that excludes affected tasks, with the raw score shown separately. Grok 4.7, Gemini Flash 3.8 and DeepSeek V4.1 Flash show differences between their raw and clean results, making the effect of benchmark leakage visible.
Use the index as a routing guide
Running every model through the same Hermes Agent harness reduces variation from model-specific runtimes and decoding setups. Reporting cost beside performance also supports practical routing decisions, including whether to use one model for every task or reserve expensive models for escalations.
The results remain specific to the Hermes Agent runtime, this task mix, high reasoning settings and single-attempt evaluation. A different framework, lower reasoning budget, repeated attempts or production-specific tools could change both performance and cost. Provider pricing and token consumption can also shift, so teams should validate shortlisted models against representative workloads.
At the listed prices, DeepSeek V4.1 Flash scores 36.91 for $0.259 per task, while GPT 6 Astra scores 56.25 for $11.61. Sonnet 5.5 gains 9.04 index points over GPT 6 Sol for another $0.59 per task. Opus 5.5 adds 10.17 points over Sonnet for another $2.17. Those increments give developers concrete thresholds for choosing a default model, defining escalation tiers and controlling agent costs at scale.