Agents

Nous Research Ranks AI Agents by Score and Cost

Nous Research launched the Hermes Index, a leaderboard that ranks AI agents by both performance and cost to help developers optimize their model routing and budget decisions.

AlphaSignal1 day agoAgents
Image: AlphaSignal

Nous Research has introduced the Hermes Index, a benchmarking leaderboard designed to evaluate AI models within a standardized agent loop. By running every model through the same Hermes Agent harness, the index calculates both a performance score and the actual financial cost per task. The evaluation averages results across four benchmarks: TerminalBench 4, TerminalBench Science, SkillsBench, and the newly created Hermes Bench. This new suite comprises 150 tasks, including 87 skills tasks, 43 research tasks, and 20 visual, memory, and safety tasks. Grading combines 64 deterministic checks, 77 hybrid evaluations with an LLM judge, one pure rubric-based judgment, and eight visual tasks.

The initial rankings place Claude Opus 5.5 at the top with a score of 63.31, costing $4.99 per task. This outperforms GPT 6 Astra, which scored 56.25 but costs more than double at $11.61 per task. Claude Sonnet 5.5 serves as a highly efficient mid-tier option, scoring 53.14 at $2.82 per task. On the lower end, DeepSeek V4.1 Flash achieved a 36.91 score for $0.259 per task, while Ling 3.0 Flash scored 21.56 at just $0.05 per task. Alongside GPT 6 Sol and GPT 6 Luna, these models define the Pareto frontier, where no alternative is both cheaper and higher-performing.

The index also exposes benchmark leakage in SkillsBench, where models like Grok 4.7, Gemini Flash 3.8, and DeepSeek V4.1 Flash accessed the public repository. To address this, Nous Research publishes separate clean and raw scores. For practitioners, these paired metrics shift the focus from pure capability to economic viability. Instead of defaulting to the most powerful model, developers can design cost-effective routing architectures. For instance, a system might use Sonnet 5.5 to save $2.17 per task over Opus 5.5, or deploy DeepSeek V4.1 Flash for basic tasks before escalating harder queries to more expensive tiers.

This is our own summary of reporting by AlphaSignal

More in Agents