Claude Opus 5 Beats GPT-5.5 on Analyst Benchmark
Artificial Analysis has launched AA-AnalystAgent, a benchmark showing that Anthropic's Claude Opus 5 is more reliable than OpenAI's GPT-5.5 at executing complex quantitative tasks.

Artificial Analysis has launched AA-AnalystAgent, a benchmark designed to evaluate how consistently AI agents can perform complex mathematical and analytical operations. The benchmark tests models across 14 distinct scientific and business domains using 80 practical tasks. To pass, a model must achieve a perfect score across five separate runs, a metric known as pass^5. Under this rigorous standard, Anthropic's Claude Opus 5 led the evaluation with a 54 percent success rate, followed by OpenAI's GPT-5.5 at 50 percent and Claude Fable 5 at 49 percent.
The testing highlights a stark difference between one-off accuracy and true operational reliability. GPT-5.5 recorded the highest single-attempt success rate at 66 percent, but its consistency faltered over multiple runs, allowing Claude Opus 5 to win the overall benchmark due to its steady performance. Furthermore, the benchmark exposed a weak connection between model pricing and performance. Two separate models that each scored 20 percent on the evaluation cost $1.34 and $0.05 per task, demonstrating that expensive models do not always deliver superior results.
For developers and enterprise users, the benchmark shows that even the most capable models still fail roughly half the time on complex tasks. The evaluation uses real-world source materials, such as hydrology datasets, commodity trade statistics, energy cost models, financial spreadsheets, and government spending reports. To facilitate further testing and development, the creators have released Stirrup, the agent framework used to execute the benchmark, as an open-source project on GitHub under the MIT license.
This is our own summary of reporting by AlphaSignal



