LIVE FEED
Anthropic and OpenAI Open Doors to Embedded Safety Evaluators

Anthropic and OpenAI Open Doors to Embedded Safety Evaluators

FIRST LOOK ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 8.2 TechCrunch AI

Anthropic and OpenAI have proposed embedding independent third-party safety evaluators — including organisations like METR and Redwood Research — directly inside frontier AI companies, granting access to training checkpoints, post-training environments, and evaluation logs rather than only finished models. This closes a critical oversight gap: defenders and policymakers have historically had no mechanism to verify whether alignment claims made by AI labs actually held during training, leaving assurance entirely self-reported. Significant implementation detail remains unresolved, including scope of access, disclosure rights, and whether the arrangement will be codified in legislation or remain voluntary.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.