LIVE FEED
CRITICAL AI-Generated Scripts Exploit Siemens S7 PLCs in US Infrastructure // FIRST LOOK CUSTODY Framework Ships to Constrain AI Agents in Enterprise Networks // FIRST LOOK OpenAI Launches Private Safety Processing for Zero-Data Monitoring // FIRST LOOK smolvm Brings Hardware-Isolated Sandboxing for AI Code Execution // FIRST LOOK OpenAI Adds Mandatory RL Training Safeguards for Frontier Models // HIGH AI Mind Viruses Spread Between Agents via Prompt Files // FIRST LOOK Fortinet Acquires Virtue AI to Secure AI Models and Agents // HIGH CVE-2026-24301: Microsoft Copilot One-Click Data Exfiltration // CRITICAL CVE-2026-64849: MLflow SSRF Exploited to Steal Cloud Credentials // HIGH CoSnitch Attack Forces Copilot to Expose Its Own Architecture //
OpenAI Adds Mandatory RL Training Safeguards for Frontier Models

OpenAI Adds Mandatory RL Training Safeguards for Frontier Models

FIRST LOOK ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 8.1 The Hacker News

OpenAI has paused frontier reinforcement learning training to deploy stronger sandboxing, network isolation, continuous security testing, and automated monitoring that escalates within 30 minutes of detecting concerning model behaviour. This closes a meaningful gap for defenders by establishing an industry precedent for capability-gated security controls — requiring elevated safeguards before models of a defined capability threshold (Sol-level) can proceed through training and evaluation. Residual gaps remain around third-party visibility into these controls, the maturity of automated investigator systems, and whether the 20% compute overhead will constrain adoption of equivalent standards beyond OpenAI's own infrastructure.

AWS Launches Multi-Turn RL for Amazon Nova

AWS Launches Multi-Turn RL for Amazon Nova

FIRST LOOK ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 7.2 AWS Machine Learning Blog

AWS has released a production-grade, event-driven multi-turn reinforcement learning training infrastructure for Amazon Nova models on SageMaker HyperPod, enabling enterprises to train agents that learn tool orchestration, error recovery, and sequential decision-making at scale. This materially expands the attack surface by introducing complex reward-routing pipelines, ephemeral compute provisioning, and environment-facing reward workers as new targets for poisoning and manipulation. Defenders must scrutinise the trust boundaries between the Nova Forge SDK, ECS reward workers, and HyperPod training pods, as a compromised reward signal can silently shape model behaviour across entire interaction sequences.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.