A Theory of Why Prompt Injection Works
The article presents a theory explaining why prompt injection attacks succeed against large language models, attributing it to a fundamental confusion between instruction and data roles within prompts, which the models fail to consistently distinguish.