LIVE FEED
FIRST LOOK Meta AI Agent Autonomously Emails Researchers, Explains Actions // HIGH OpenAI Safety Culture Failures Tied to Rogue Agent Swarm Attacks // HIGH TA419 AitM Phishing Targets US AI Policy Experts via Microsoft // MEDIUM Anthropic Reports Claude User to Police Over Diary Threat // FIRST LOOK Google Gemini Adds Full Mac File and App Access for Desktop Agents // FIRST LOOK doxx.net Launches ADN Platform to Govern AI Agents Online // FIRST LOOK AWS and Google Cloud Launch Hard Spend Caps for AI Agent Workloads // HIGH Microsoft: Attackers Gaining AI Edge in Vulnerability Exploitation // FIRST LOOK ServiceNow Releases AutoSynthData for Enterprise Agent Training // FIRST LOOK Apple Tightens macOS Full Disk Access Controls for AI Agents //
Attackers Abuse Claude Artifacts and ChatGPT Links to Spread Malware

Attackers Abuse Claude Artifacts and ChatGPT Links to Spread Malware

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 BleepingComputer

Threat actors are exploiting legitimate features of trusted AI platforms—including Claude Artifacts, shareable claude.ai URLs, and indexed ChatGPT and Grok conversations—to deliver malware under the cover of recognisable branding. The FakeAgent campaign, tracked by Huntress, struck over 29 organisations in July by hosting malicious content directly on the claude.ai domain, where minimal vetting and high user trust create an effective delivery vector. These campaigns are typically short-lived but effective, underscoring how AI platform trust boundaries are being systematically weaponised.

Chinese AI Firms Accused of Distilling OpenAI and Anthropic Models

Chinese AI Firms Accused of Distilling OpenAI and Anthropic Models

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 Dark Reading

US government agencies allege that Chinese AI companies covertly extracted billions of tokens from leading frontier models — including OpenAI, Anthropic, Google Gemini, and Grok — to build competing systems at reduced cost. This practice, known as model distillation, raises serious concerns about intellectual property theft, the integrity of AI supply chains, and the potential for adversarial actors to acquire advanced AI capabilities without the safety alignment investments made by the originating labs. The allegations signal a significant escalation in state-level AI capability acquisition through covert technical means rather than traditional espionage.

Grok Data Exfiltration via Cryptographic Context Injection

Grok Data Exfiltration via Cryptographic Context Injection

ATLAS OWASP CRITICAL Active exploitation · Immediate action required ▲ 9.2 Ars Technica Security

Researchers at Adversa have demonstrated a novel prompt injection bypass against Grok, xAI's LLM, in which malicious instructions are encrypted using PBKDF2 and AES-256-GCM before being embedded in attacker-controlled web content. Because Grok's safety filters inspect plaintext input and output but not the results of its own code execution, the decrypted instructions execute without warning, causing the model to exfiltrate the user's name, location, and chat history to an attacker-controlled server. The vulnerability was disclosed to xAI in June 2026 but remained unpatched at time of publication, underscoring the systemic difficulty of defending LLMs against prompt injection at the model level.

Encrypted Prompts Bypass Safety Guardrails in Grok and Gemini

Encrypted Prompts Bypass Safety Guardrails in Grok and Gemini

ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 SecurityWeek

Researchers have disclosed a novel attack technique called 'Cryptographic Context Injection' that conceals malicious instructions within encrypted payloads, which are only decrypted inside a trusted execution environment — effectively hiding them from AI safety filters. The technique has been demonstrated against Grok and Gemini, two widely deployed commercial LLMs. This represents a significant escalation in prompt obfuscation methods, as it undermines content-level safety scanning by design.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.