Defender Impact
This research provides a mechanistic explanation for one of AI evaluation’s most persistent confounds: LLM self-reports about capabilities, limitations, and safety posture are partially determined by deployment infrastructure, not solely by trained weights. For defenders relying on model-stated boundaries as part of a safety case, this finding materially changes the evidentiary weight those statements should carry.
Capability Overview
Published by Jędrzej Maczan and accepted to COLM 2026 and KONVENS 2026, this paper demonstrates that the chat template applied during inference functions as a near-binary switch over LLM self-referential language. Across eight popular open-source instruct models up to 9B parameters, the presence of a chat template reliably increases disclaimer language (‘I’m just a language model’) and suppresses experiential language (‘I feel’, ‘I believe’). Remove the template, and the pattern inverts.
More importantly for defenders, the researchers identify a specific direction in activation space within three tested models that mechanistically reproduces this behaviour. Adding this direction to a model without a chat template causes it to disclaim as though the template were present; removing it from a model with a template suppresses disclaimers. A randomly chosen direction of equivalent magnitude produces negligible effect, confirming the direction is behaviourally meaningful rather than a noise artefact.
The practical implication is significant: what a model says about itself is a function of how it is deployed, not only what it learned. Two deployments of the same model, differing only in chat template, will produce systematically different self-descriptions — and those differences are internally representable and steerable.
Defensive Advances
Evaluators can now control for deployment context in capability assessments. Previously, comparing self-reports across models or deployments was confounded by unknown template differences. This work gives evaluators a principled variable to hold constant or vary deliberately.
Activation-steering probes provide an internal audit channel. Rather than inferring model posture purely from outputs, security teams with interpretability tooling can probe whether a given deployment’s disclaimer behaviour is template-driven or weight-driven — a meaningful distinction when assessing whether a safety-relevant behaviour is robust or surface-level.
Red teams gain a reproducible test for deployment wrapper influence. If an operator claims their deployed model ‘always refuses’ or ‘always discloses limitations’, this methodology allows testers to determine whether that behaviour is intrinsic or a chat-template artefact that could be altered by changing the deployment configuration.
Residual Gaps
The study is currently bounded to models up to 9B parameters. Whether the same activation direction generalises to frontier-scale models (70B+, or closed-weight commercial models) is an open question. Defenders should treat the methodology as validated at smaller open-source scale pending replication at larger sizes.
Operationalising activation-steering audits requires white-box model access and interpretability infrastructure that most enterprise security teams do not currently operate. The gap between ’this direction exists’ and ‘we can audit this in production’ is a meaningful maturity step. Organisations without existing mechanistic interpretability capability will need to either build tooling or rely on third-party evaluation providers.
The paper also does not address multimodal models or instruction-tuned models trained with reinforcement learning from human feedback at scale, where the relationship between template and activation may differ.
Framework Mapping
- AML.T0063 (Discover AI Model Outputs) and AML.T0069 (Discover LLM System Information): This research directly supports defenders building a more accurate picture of model output provenance by distinguishing template-conditioned from weight-conditioned behaviour.
- AML.T0056 (LLM Meta Prompt Extraction): Understanding how chat templates shape self-referential responses improves defenders’ ability to assess what system-level information is being surfaced or suppressed.
- LLM09 (Overreliance): The most direct OWASP mapping — treating LLM self-reports as ground truth is a documented overreliance risk this work concretely substantiates and provides tools to mitigate.
Deployment Considerations
Organisations should sequence adoption in three stages. First, update evaluation policy to treat LLM self-descriptions as deployment-context-dependent claims requiring corroboration. Second, for open-source model deployments, build or adopt baseline evaluation pipelines that test model behaviour both with and without the operational chat template. Third, for teams with interpretability capability, the activation direction methodology described in the paper offers a more robust internal audit path.
Complementary controls include red-teaming chat template variations as part of deployment change management, and documenting which self-referential behaviours are template-conditioned versus weight-conditioned in model cards for internal governance purposes.
Defender Checklist
- Update AI evaluation policy: LLM self-reports are deployment artefacts, not intrinsic model facts
- Add chat-template-absent baseline to standard LLM evaluation pipelines
- Document chat template configuration in deployment change management records
- Review any safety cases that cite model self-descriptions as evidence of capability limits
- Assess internal interpretability tooling maturity for activation-steering audits
- Monitor for replication of findings at larger parameter scales
References
- Maczan, J. (2026). ‘As a Language Model…’: Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It. arXiv:2609.25021. https://arxiv.org/abs/2609.25021