Research

MLCommons Outlines Rules to Spot Fake AI Benchmark Scores

Non-profit consortium MLCommons has issued a guide to combat 'benchmark washing,' helping enterprise buyers identify rigged or contaminated AI performance claims.

ML Commons21 hrs agoResearch
Image: ML Commons

AI benchmarking consortium MLCommons has published a guide warning enterprise buyers against 'benchmark washing,' where vendors selectively present convenient evaluation results to imply unearned model readiness. To help practitioners distinguish between research tools and industrialized benchmarks, the organization outlined key integrity controls. The warning comes as public leaderboards face severe contamination, where models memorize test data rather than demonstrating actual capability.

The scale of this issue is evident in recent evaluations. On a standard grade-school arithmetic benchmark, model accuracy plunged by 13 points when tested on a fresh, uncontaminated dataset. Similarly, on the public SWE-Bench Verified, top-tier models frequently score above 70%. However, when Scale AI tested models on its private SWE-Bench Pro repository, GPT-5 plummeted from 23% to under 15%, and Claude Opus 4.1 dropped from 23% to 18%. This represents a massive 55-point discrepancy between public and private testing environments.

To address these gaps, MLCommons advocates for rigorous, industrialized benchmarks like its own MLPerf Inference, currently on version 6.1, which enforces strict rules on non-determinism and replicability. The consortium also launched AILuminate in 2024 to measure safety and jailbreak resilience, complementing performance metrics. Other frameworks like Stanford's BetterBench, which evaluates 24 benchmarks against 46 criteria, and HELM also push for multi-metric evaluations. Meanwhile, studies like 'The Leaderboard Illusion' by Singh et al. in 2025 and a 2025 European Commission Joint Research Centre meta-review of roughly 100 studies highlight how easily public leaderboards are gamed.

For enterprise practitioners, these findings change how AI procurement and risk management must be handled. Teams can no longer rely on a single accuracy score to justify high-stakes investments. Instead, they must demand audited, reproducible evidence trails, check for test-set leakage, and ensure that automated LLM judges are anchored to human standards. Following structured approaches like the NIST AI Risk Management Framework of 2023 will help practitioners align their evaluation rigor with the actual consequences of model failure.

This is our own summary of reporting by ML Commons

More in Research