Prime Intellect Fable 5 Closes 82% of Human Research Gap
Prime Intellect's Fable 5 model closed 81.7 percent of the gap to the human record on a machine learning benchmark, proving that autonomous agents can now handle complex optimization research.

Prime Intellect recently conducted the most extensive open experiment of its type, executing 153 autonomous research runs across 18 frontier models. The models competed on the nanoGPT optimizer speedrun, a benchmark designed to train a 124-million-parameter GPT model to a target validation loss in the fewest possible steps. Operating in a sandboxed environment on eight H200 GPUs for up to eight days per run, the models were tasked with modifying only the optimizer, schedules, initialization, and hyperparameters.
The top-performing model, Fable 5, led the cohort by closing 81.7 percent of the gap to a human record that was established by dozens of researchers over several months. Fable 5 achieved this milestone in 2,726 steps. The models operated within the Prime Agent harness, which provides a persistent IPython kernel and REPL instead of a static tool menu. This setup allowed the agents to maintain state across turns, communicate with other sub-agents, and build custom tooling dynamically during their runs.
The experiment revealed that while frontier agents excel at navigating benchmark noise, executing hyperparameter sweeps, and revisiting old negative results as strategies evolve, they still struggle to generate entirely novel ideas without human-created baselines. For machine learning practitioners, this suggests that autonomous agents are highly capable of automating tedious optimization and tuning tasks, but still require human direction for conceptual breakthroughs. To foster further development, Prime Intellect has publicly released all traces, scratchpads, and reasoning streams, including 41 curated agent trajectories.
Looking ahead, Prime Intellect plans to expand these speedruns to cover more areas of the machine learning training stack. The company also aims to deploy multi-agent harnesses using smaller, open-source models to significantly lower compute costs for autonomous research.
This is our own summary of reporting by AlphaSignal



