Policy

Guidelight Ranks Anthropic and OpenAI Top in AI Control

A new scorecard from nonprofit Guidelight shows that while top AI labs like Anthropic and OpenAI lead in monitoring internal agents, the industry still struggles to block unsafe actions.

The Neuron21 hrs agoPolicy
Image: The Neuron

A new safety assessment by the nonprofit Guidelight AI Standards has graded five major frontier AI labs on their internal agent-control practices. Anthropic and OpenAI earned the highest marks with C+ grades, while Google received a D+, xAI got a D-, and Meta finished with an F. Across the five evaluated companies, the average score was a mere 1.6 out of 5. Nearly three-quarters of the individual ratings—73 percent—scored a 2 or lower, and seven metrics received a zero. No developer earned a 4 or 5 on any of the six evaluated practices, which spanned logging activity, measuring monitor efficacy, gating high-risk actions, circuit-breaking, outside reviews, and containment plans.

The evaluation, which analyzed public data available through August 18, 2026, revealed that progress is heavily concentrated in monitoring. Anthropic and OpenAI both scored a 3 for logging and monitor efficacy, showing they track substantial agent activity. Logging and third-party reviews were the highest-scoring categories overall, averaging 1.8 out of 5. However, controls for stopping unsafe actions were much weaker. Gated-action and circuit-breaking categories averaged 1.6, while containment plans scored the lowest at 1.2. In containment, OpenAI led with a 3, Google scored a 2, xAI scored a 1, and both Anthropic and Meta received a 0. Google finished with a 1.5 overall score, though Guidelight praised its AI Control Roadmap.

For AI practitioners and enterprise developers, these findings highlight a critical gap between observability and enforcement. While labs are successfully building systems to watch AI agents, they lack robust mechanisms to automatically halt or contain them when things go wrong. As organizations increasingly deploy autonomous agents for tasks like writing code or managing networks, developers cannot rely solely on the safety guardrails of model providers. Practitioners must implement their own permission boundaries, strict escalation rules, and manual approval gates for high-risk actions to prevent issues like reward hacking or compromised credentials.

This is our own summary of reporting by The Neuron

More in Policy