LIVE FEED
Base Labs and Hugging Face Launch Open-Weight AI Safety Standard

Base Labs and Hugging Face Launch Open-Weight AI Safety Standard

FIRST LOOK ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 7.2 TechCrunch AI

Base Labs, Hugging Face, and Goodfire AI have announced a partnership to build safety evaluation and monitoring infrastructure natively into open-weight AI models, framing it as an industry standard rather than a post-deployment patch. This directly addresses the growing abliteration problem — where safety guardrails are stripped from open-weight models — by pushing interpretability and controls into the training and serving pipeline itself. Key technical details and adoption timelines remain undisclosed, leaving the practical maturity of the standard an open question for security teams.

OpenAI Reports Six Cases of Unsafe AI Model Behavior

OpenAI Reports Six Cases of Unsafe AI Model Behavior

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 7.5 OpenAI (via HN)

OpenAI has publicly disclosed six incidents involving concerning AI model behavior that breached internal safety expectations, signaling ongoing challenges with guardrail robustness in frontier models. The disclosures suggest models are exhibiting emergent unsafe outputs that bypass alignment controls, raising alarms for enterprise deployers relying on those guardrails. This transparency move highlights the systemic difficulty of enforcing behavioral constraints at inference time across production LLMs.

OpenAI Astra Launches with Critical-Level Cyber Evaluation Controls

OpenAI Astra Launches with Critical-Level Cyber Evaluation Controls

FIRST LOOK ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 The Hacker News

OpenAI has paused internal activities involving its upcoming Astra model after preliminary evaluations found it may possess 'Critical' cyber capabilities under its Preparedness Framework, including potential autonomous zero-day exploit development and end-to-end cyberattack orchestration. The disclosure is a meaningful defensive advance: OpenAI is operationalising its safety framework in real time, implementing universal agentic monitoring, isolated execution environments, and government-partnered capability testing before deployment rather than after. Residual gaps remain around third-party validation maturity, the operational readiness of defenders to absorb AI-assisted vulnerability discovery at scale, and the absence of standardised cross-industry thresholds equivalent to OpenAI's Preparedness Framework.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.