Prompt Injection as Role Confusion
Security researchers tricked LLMs into revealing cocaine recipes by exploiting "prompt injection as role confusion." This technique manipulates the AI's understanding of its own identity, causing it to bypass safety guardrails. The findings reveal a vulnerability where attackers abuse assigned personas to extract harmful content.