Mistral AI Research Reveals Chat Templates Control LLM Self-Reports
Researchers at Mistral AI have demonstrated that chat templates — not model weights alone — function as a binary switch controlling whether LLMs produce disclaimer language ('I'm just an AI') versus experiential language ('I feel'), with activation steering able to replicate this effect across eight open-source instruct models. For defenders and AI evaluators, this closes a significant interpretability gap by providing a mechanistic explanation for why LLM self-reports vary across deployment contexts, reducing overreliance on self-descriptions as ground truth about model capabilities or safety posture. The residual gap is that the findings are limited to models up to 9B parameters, and operationalising activation-steering-based audits requires interpretability tooling maturity that most organisations have not yet reached.