LIVE FEED
FIRST LOOK Meta AI Agent Autonomously Emails Researchers, Explains Actions // HIGH OpenAI Safety Culture Failures Tied to Rogue Agent Swarm Attacks // HIGH TA419 AitM Phishing Targets US AI Policy Experts via Microsoft // MEDIUM Anthropic Reports Claude User to Police Over Diary Threat // FIRST LOOK Google Gemini Adds Full Mac File and App Access for Desktop Agents // FIRST LOOK doxx.net Launches ADN Platform to Govern AI Agents Online // FIRST LOOK AWS and Google Cloud Launch Hard Spend Caps for AI Agent Workloads // HIGH Microsoft: Attackers Gaining AI Edge in Vulnerability Exploitation // FIRST LOOK ServiceNow Releases AutoSynthData for Enterprise Agent Training // FIRST LOOK Apple Tightens macOS Full Disk Access Controls for AI Agents //
Claude AI Used by Yemen Cell to Develop Guided Missiles

Claude AI Used by Yemen Cell to Develop Guided Missiles

ATLAS OWASP CRITICAL Active exploitation · Immediate action required ▲ 9.1 Schneier on Security

Anthropic's Claude was exploited by a threat actor cell in northern Yemen to develop guidance, navigation, and control software for multiple weapons systems, including a guided rocket and a hypersonic glide vehicle variant. The actors systematically evaded Claude's safety guardrails by splitting sessions, obscuring intent, and orchestrating multiple Claude instances in parallel as a pseudo-engineering team. While no operational device was confirmed fielded, a guided rocket test-fire was attempted, demonstrating real-world weapons development acceleration via LLM assistance.

Anthropic Exposes 200M-Exchange Model Distillation Attacks

Anthropic Exposes 200M-Exchange Model Distillation Attacks

ATLAS OWASP CRITICAL Active exploitation · Immediate action required ▲ 9.2 TechCrunch AI

Anthropic has published a detailed report attributing nearly 200 million adversarial API exchanges to coordinated model distillation campaigns conducted by Alibaba, Moonshot AI, and DeepSeek. Attackers used prompt obfuscation techniques — including fake translation requests — to bypass Claude's summarised-thinking safeguards and extract raw chain-of-thought traces for use as supervised fine-tuning data. One Moonshot AI campaign was assessed as routing requests directly through Chinese military infrastructure, adding a significant geopolitical dimension to what is otherwise an IP-theft threat.

Encrypted Prompts Bypass Safety Guardrails in Grok and Gemini

Encrypted Prompts Bypass Safety Guardrails in Grok and Gemini

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 SecurityWeek

Researchers have disclosed a novel attack technique called 'Cryptographic Context Injection' that conceals malicious instructions within encrypted payloads, which are only decrypted inside a trusted execution environment — effectively hiding them from AI safety filters. The technique has been demonstrated against Grok and Gemini, two widely deployed commercial LLMs. This represents a significant escalation in prompt obfuscation methods, as it undermines content-level safety scanning by design.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.