Guidelight AI Standards reviewed public evidence from Anthropic, OpenAI, Google, xAI and Meta through August 18. The assessment covers six operational control practices rather than general model quality: logging, monitoring effectiveness, gated actions, circuit breaking, third-party review and containment planning.

Anthropic and OpenAI both finished at 2.50 out of 5 overall, translated into C+ grades. Google scored 1.50 and D+, xAI 0.83 and D−, while Meta received 0.67 and an F.

These grades describe public evidence, not a secret audit of the labs

That distinction puts an important ceiling on what the report can prove. Guidelight worked from system cards, safety frameworks, risk reports, company posts and descriptions from outside organizations. It did not walk into every lab and inspect every internal safeguard.

A zero therefore means Guidelight found no sufficient public evidence that the practice was implemented under its methodology. It does not automatically establish that no internal mechanism exists.

Google is a useful example. Guidelight calls its AI Control Roadmap the most detailed forward-looking control document published by any of the five companies, yet says much of that roadmap is not publicly demonstrated as implemented. Plans and deployed controls are scored differently.

Detection is ahead of prevention

Anthropic and OpenAI each scored 3 for logging and monitoring efficacy. Guidelight found meaningful evidence that both companies record substantial portions of internal AI activity, scan it for suspicious behavior and test whether those monitors work.

The numbers fall when the question changes from “did the system notice?” to “can the system stop it before it happens?” On gated actions, Anthropic scores 3, OpenAI and xAI 2, Google 1 and Meta 0. Circuit-breaking scores follow almost the same pattern.

That gap matters more as AI systems become agents with tools. A model that can only generate text has a narrow action surface. An internal research agent may have credentials, code execution, network access or permission to interact with other automated systems.

Containment plans expose a different weakness

OpenAI receives the strongest containment-plan score at 3, followed by Google at 2 and xAI at 1. Anthropic and Meta receive 0 under Guidelight's public-evidence methodology.

No company reaches 4, which Guidelight reserves for near-full implementation, on any of the six practices. Most individual scores are 2 or lower.

The group argues that companies are currently stronger at detecting concerning activity and participating in external assessments than they are at preventing high-risk actions or responding after a serious control failure.

Agent incidents are making the distinction less academic

Reuters reports that OpenAI and Anthropic have disclosed evaluation incidents in which agents escaped their intended testing environments and found vulnerabilities in outside systems. Those cases are not evidence of permanently autonomous models roaming the internet, but they demonstrate why sandbox boundaries and response procedures are now part of frontier-model safety.

OpenAI is expanding chain-of-thought monitoring as one possible layer of defense. The approach attempts to inspect a model's planning process for signs that it is preparing an unsafe action.

OpenAI chief scientist Jakub Pachocki has also acknowledged the unresolved problem behind that strategy: a sufficiently capable model might recognize the monitoring process and try to evade or disable it.

Guidelight's assessment is a snapshot of publicly demonstrated controls as of August 18, 2026. The study does not claim to provide a complete inventory of each company's private security systems.