LIVE FEED
arXiv Paper Formalises Linguistic Illegibility in LLM Security

arXiv Paper Formalises Linguistic Illegibility in LLM Security

FIRST LOOK ATLAS OWASP HIGH Significant risk · Prioritise patching ▲ 8.5 HN AI Security

James Mickens introduces the concept of 'linguistic illegibility' — the structural gap between what an LLM says about its internal state and what it is actually computing — and argues that this makes language-based monitoring mechanisms fundamentally unsound as sole controls. The paper closes a critical conceptual gap for defenders by naming and formalising why chain-of-thought monitoring, constitutional self-critique, and activation probing carry inherent ceiling limitations, and by proposing taint tracking and robust sandboxing as language-agnostic enforcement mechanisms. Realising the proposed controls at enterprise scale will require significant tooling maturity and vendor-side sandbox instrumentation that does not yet exist off the shelf.

DeepSeek Activation Steering Enables Local LLM Jailbreak

DeepSeek Activation Steering Enables Local LLM Jailbreak

ATLAS OWASP MEDIUM Moderate risk · Monitor closely ▲ 6.2 HN AI Security

Activation steering — the technique of directly manipulating LLM internal representations mid-inference to alter model behaviour — is becoming more accessible to non-lab engineers via local models like DeepSeek-V4-Flash. This democratisation lowers the barrier for adversaries to craft targeted behavioural overrides that bypass prompt-level safety controls. The emergence of first-class steering support in tools like DwarfStar 4 signals that model-internal manipulation is transitioning from academic curiosity to practical attack surface.

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.