The Hallway Track
Research Findings

Prompt Injection as Role Confusion

Simon Willison · Jun 22, 2026 · Research Findings

Models confuse text style with role, making prompt injection a 'role confusion' problem defenses can't easily fix.

“destyling causes average attack success in our dataset to plunge from 61% to 10%”

Researchers Charles Ye, Jasmine Cui, and Dylan Hadfield-Menell show LLMs distinguish privileged from untrusted text by writing style rather than content, a mechanism they call 'role confusion' that enables reliable jailbreaks. Rewriting injected text to look less like a model's internal formatting ('destyling') cut attack success from 61% to 10%, suggesting injection defense will remain whack-a-mole until models achieve genuine role perception.

prompt-injection jailbreaking llm-security role-confusion

Watch / read the original source →