LIVE FEED
PuzzleMask Bypasses LLM Policy Guards Using Plain Prose

PuzzleMask Bypasses LLM Policy Guards Using Plain Prose

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 Check Point Research

Check Point Research has disclosed PuzzleMask, a prompt-crafting technique that embeds policy-violating payloads inside ordinary English prose to fool lightweight LLM-based gatekeepers into classifying malicious input as benign. Tested against four commercial and open-source safety models, the technique achieved a 100% bypass rate on gatekeeper checks, with the downstream target model successfully extracting and acting on the hidden payload in over 90% of trials. The attack requires no special encoding, invisible characters, or emoji obfuscation, making it harder to detect with traditional content filters.

CoSnitch Attack Forces Copilot to Expose Its Own Architecture

CoSnitch Attack Forces Copilot to Expose Its Own Architecture

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.2 Dark Reading

Researchers demonstrated a 'meta-hacking' technique dubbed CoSnitch that manipulates Microsoft Copilot into disclosing its own internal security weaknesses and architectural details. The attack leverages the AI system's own reasoning capabilities against itself, effectively turning the assistant into an unwitting reconnaissance tool. This class of vulnerability has significant implications for enterprise deployments where Copilot has access to sensitive organisational infrastructure and data.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.