LIVE FEED
Mistral AI Research Reveals Chat Templates Control LLM Self-Reports

Mistral AI Research Reveals Chat Templates Control LLM Self-Reports

FIRST LOOK ATLAS OWASP LOW Limited impact · Standard review ▲ 5.8 Mistral AI (via HN)

Researchers at Mistral AI have demonstrated that chat templates — not model weights alone — function as a binary switch controlling whether LLMs produce disclaimer language ('I'm just an AI') versus experiential language ('I feel'), with activation steering able to replicate this effect across eight open-source instruct models. For defenders and AI evaluators, this closes a significant interpretability gap by providing a mechanistic explanation for why LLM self-reports vary across deployment contexts, reducing overreliance on self-descriptions as ground truth about model capabilities or safety posture. The residual gap is that the findings are limited to models up to 9B parameters, and operationalising activation-steering-based audits requires interpretability tooling maturity that most organisations have not yet reached.

Base Labs and Hugging Face Launch Open-Weight AI Safety Standard

Base Labs and Hugging Face Launch Open-Weight AI Safety Standard

FIRST LOOK ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 7.2 TechCrunch AI

Base Labs, Hugging Face, and Goodfire AI have announced a partnership to build safety evaluation and monitoring infrastructure natively into open-weight AI models, framing it as an industry standard rather than a post-deployment patch. This directly addresses the growing abliteration problem — where safety guardrails are stripped from open-weight models — by pushing interpretability and controls into the training and serving pipeline itself. Key technical details and adoption timelines remain undisclosed, leaving the practical maturity of the standard an open question for security teams.

OpenAI Astra Ships Recurrent Depth Reasoning with CoT Monitoring Pledge

OpenAI Astra Ships Recurrent Depth Reasoning with CoT Monitoring Pledge

FIRST LOOK ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 7.2 TechCrunch AI

OpenAI's Astra model introduces 'recurrent depth' (opaque recurrence), a non-linear reasoning technique that processes queries in iterative loops rather than sequential chain-of-thought steps. The development is significant for defenders because it tests the limits of chain-of-thought monitoring — a primary mechanism for detecting AI misalignment and rogue agent behaviour — while OpenAI's accompanying commitment to legible CoT and structured monitoring programs provides a concrete defensive baseline to evaluate against. Residual gaps centre on the absence of standardised monitorability requirements across labs, the immaturity of interpretability tooling for looped inference, and the risk that competitive pressure could erode the CoT-faithfulness norms that currently underpin AI oversight.

DeepSeek Activation Steering Enables Local LLM Jailbreak

DeepSeek Activation Steering Enables Local LLM Jailbreak

ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 6.2 HN AI Security

Activation steering — the technique of directly manipulating LLM internal representations mid-inference to alter model behaviour — is becoming more accessible to non-lab engineers via local models like DeepSeek-V4-Flash. This democratisation lowers the barrier for adversaries to craft targeted behavioural overrides that bypass prompt-level safety controls. The emergence of first-class steering support in tools like DwarfStar 4 signals that model-internal manipulation is transitioning from academic curiosity to practical attack surface.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.