Claude Now Leads 26% of Anthropic's AI Research
Anthropic has proposed three transparency metrics for frontier labs after revealing that its Claude model now autonomously directs over a quarter of the company's own AI research.

Anthropic has introduced a trio of reporting metrics designed to help frontier artificial intelligence laboratories track and share their progress in automating research. Alongside the proposal, the company disclosed internal data from July and August 2026 showing that its Claude model now leads 26% of Anthropic's internal AI research and development tasks. This represents a massive surge from less than 1% earlier in the year. According to the Epoch AI scale, more than 90% of Anthropic's R&D work has reached the "AI collaborates" level (AL3) or higher, with Claude driving 26% of tasks at the "AI leads" level (AL4). No tasks have yet reached the fully autonomous AL5 level, though the company projects AL4 tasks could reach 80% by the end of 2026.
To generate these metrics, a Claude research agent analyzed internal Slack messages and documents to categorize roughly 15,000 tasks into 542 nodes, which a separate Claude judge then graded. Beyond research automation, Anthropic's second metric tracks agent oversight. The firm reported that approximately 30,000 internal agents run concurrently on its primary engineering platform. These agents generated over one billion decisions in August 2026, with online monitors blocking 0.002% of actions, or about one in every 47,000. Meanwhile, offline monitors flag roughly 100,000 transcripts weekly, resulting in about 50 escalations to human reviewers.
The final metric measures compute resources. During a single week in July 2026, Anthropic dedicated 6% of its overall AI R&D compute to safety research, a figure that rose to 12% when looking strictly at AI-driven R&D compute. For industry practitioners, these metrics provide a concrete framework to gauge how rapidly AI-assisted development is compressing engineering timelines. As models increasingly write their own code and run their own experiments, developers will need to adapt to faster release cadences, more frequent API migrations, and accelerated evaluation schedules, all while implementing the multi-layered monitoring systems necessary to keep autonomous agents secure.
This is our own summary of reporting by AlphaSignal



