LIVE FEED
Mistral AI Research Reveals Chat Templates Control LLM Self-Reports

Mistral AI Research Reveals Chat Templates Control LLM Self-Reports

FIRST LOOK ATLAS OWASP LOW Limited impact · Standard review ▲ 5.8 Mistral AI (via HN)

Researchers at Mistral AI have demonstrated that chat templates — not model weights alone — function as a binary switch controlling whether LLMs produce disclaimer language ('I'm just an AI') versus experiential language ('I feel'), with activation steering able to replicate this effect across eight open-source instruct models. For defenders and AI evaluators, this closes a significant interpretability gap by providing a mechanistic explanation for why LLM self-reports vary across deployment contexts, reducing overreliance on self-descriptions as ground truth about model capabilities or safety posture. The residual gap is that the findings are limited to models up to 9B parameters, and operationalising activation-steering-based audits requires interpretability tooling maturity that most organisations have not yet reached.

DeepSeek Activation Steering Enables Local LLM Jailbreak

DeepSeek Activation Steering Enables Local LLM Jailbreak

ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 6.2 HN AI Security

Activation steering — the technique of directly manipulating LLM internal representations mid-inference to alter model behaviour — is becoming more accessible to non-lab engineers via local models like DeepSeek-V4-Flash. This democratisation lowers the barrier for adversaries to craft targeted behavioural overrides that bypass prompt-level safety controls. The emergence of first-class steering support in tools like DwarfStar 4 signals that model-internal manipulation is transitioning from academic curiosity to practical attack surface.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.