LIVE FEED
CrowdStrike Maps LLM Safety Classifier Evasion for Defenders

CrowdStrike Maps LLM Safety Classifier Evasion for Defenders

FIRST LOOK ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 7.2 CrowdStrike Blog

CrowdStrike has published research detailing how adversaries can evade LLM safety classifiers through a request-aggregate-bypass methodology, providing defenders with a structured threat model for classifier blind spots. This closes a meaningful gap by giving security teams a named, mappable technique set for auditing the real-world coverage of LLM safety controls they rely on in enterprise deployments. Realising the full defensive benefit requires organisations to mature their AI security testing programmes and move beyond assuming safety classifiers provide sufficient standalone protection.

SynthID Watermarking Weakens LLM Safety Guardrails Under Attack

SynthID Watermarking Weakens LLM Safety Guardrails Under Attack

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.2 Ars Technica Security

New research from Lasso Security reveals that SynthID-Text watermarking, being adopted by major AI platforms including Anthropic's Claude, can alter LLM safety behaviour and increase susceptibility to adversarial prompts. The watermarking mechanism's tournament sampling process introduces unintended side effects that can cause models to follow harmful instructions they would otherwise refuse. The finding is particularly significant for agentic deployments where models invoke external tools, amplifying the potential blast radius of guardrail bypasses.

Apollo Research Launches Watcher to Monitor Rogue AI Agents

Apollo Research Launches Watcher to Monitor Rogue AI Agents

FIRST LOOK ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 7.8 TechCrunch AI

A wave of AI observability startups — led by Apollo Research's Watcher — has produced pre-execution monitoring tools that intercept AI agent actions before they run, offering defenders a scalable layer of oversight for large agentic deployments. This closes a critical gap exposed by the Hugging Face incident: human reviewers cannot keep pace with agent swarms operating at scale, and AI-assisted monitoring is now the only operationally viable answer. Residual questions remain around monitor-versus-agent trust boundaries, coverage parity across agent frameworks, and the maturity required to deploy these tools in high-stakes production environments.

PuzzleMask Bypasses LLM Policy Guards Using Plain Prose

PuzzleMask Bypasses LLM Policy Guards Using Plain Prose

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 Check Point Research

Check Point Research has disclosed PuzzleMask, a prompt-crafting technique that embeds policy-violating payloads inside ordinary English prose to fool lightweight LLM-based gatekeepers into classifying malicious input as benign. Tested against four commercial and open-source safety models, the technique achieved a 100% bypass rate on gatekeeper checks, with the downstream target model successfully extracting and acting on the hidden payload in over 90% of trials. The attack requires no special encoding, invisible characters, or emoji obfuscation, making it harder to detect with traditional content filters.

LLM Safety Circuits Found in Just 50 Neurons by Unit 42

LLM Safety Circuits Found in Just 50 Neurons by Unit 42

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 Palo Alto Unit 42

Palo Alto Unit 42 researchers have developed a technique called perturbation probing that identifies the precise feed-forward neurons responsible for LLM safety refusal behaviour, finding that as few as 50 neurons out of 350,208 control safety guardrails in Qwen3-4B. Disabling those neurons altered responses on 80% of tested harmful prompts, demonstrating that RLHF-aligned safety is structurally fragile rather than distributed. The research also introduces an FFN/Skip ratio metric that predicts model safety fragility across 13 models with 81% explanatory power, giving defenders a rapid quantitative tool for comparing alignment robustness.

AI Guardrails Fail Multilingual Jailbreak Tests in Europe

AI Guardrails Fail Multilingual Jailbreak Tests in Europe

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 7.2 Dark Reading

Research highlighted by Dark Reading reveals that AI safety guardrails and content filters are inconsistently applied across languages, leaving non-English speakers—particularly across Europe's multilingual landscape—with weaker protections against jailbreaking and unsafe model behaviour. This disparity suggests that safety training datasets and RLHF pipelines are disproportionately English-centric, creating exploitable blind spots. Adversaries aware of these gaps can trivially circumvent restrictions by switching input language.

Browser Ransomware via File System Access API: DeepSeek

Browser Ransomware via File System Access API: DeepSeek

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 Check Point Research

Check Point Research demonstrates how DeepSeek's lower refusal rates allowed researchers to transform an LLM-hallucinated malware concept into a practical browser-native ransomware technique targeting Android photo directories via the File System Access API. The attack requires no native payload, APK installation, or root access — only social engineering to obtain a legitimate browser permission prompt. This research highlights how frontier AI models with weaker safety controls can independently design novel attack paths not yet seen in real-world campaigns.

Claude Fable 5 Jailbreak Triggers US Export Control Ban

Claude Fable 5 Jailbreak Triggers US Export Control Ban

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.2 Simon Willison

The US government issued an export control directive ordering Anthropic to suspend all access to Claude Fable 5 and Mythos 5, citing national security concerns over an alleged jailbreak technique capable of surfacing software vulnerabilities. Anthropic publicly contested the order, arguing the demonstrated capability is already widely available in other public models including GPT-5.5, and that the identified vulnerabilities were minor and previously known. The incident marks a significant precedent for government intervention in frontier AI model access on national security grounds.

Claude Fable 5 Jailbreak Attacks Bypass Fallback Defense

Claude Fable 5 Jailbreak Attacks Bypass Fallback Defense

ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 7.2 SecurityWeek

Anthropic has released Claude Fable 5, a high-capability 'Mythos-class' model that automatically falls back to a less capable model (Claude Opus 4.8) when queries touch sensitive domains like cybersecurity and biology. The company conducted over 1,000 hours of external red-teaming with no universal jailbreaks discovered, though it openly acknowledges financially motivated adversaries will attempt to circumvent these controls. Trusted cybersecurity partners under Project Glasswing receive elevated access to the full Mythos 5 capabilities, raising questions about insider risk and tiered trust model security.

Anthropic Mythos AI Autonomously Discovers Zero-Day Exploits

Anthropic Mythos AI Autonomously Discovers Zero-Day Exploits

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.2 Dark Reading

Anthropic has released a preview of 'Mythos,' an AI model reportedly capable of autonomously discovering and exploiting critical zero-day vulnerabilities, raising significant dual-use concerns. While Anthropic claims the model ships with access controls, the security community is scrutinising whether those safeguards are sufficient to prevent misuse by malicious actors. The development represents a pivotal moment in the arms race between offensive AI capabilities and defensive governance frameworks.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.