I found this out the embarrassing way by " _attacking my own system _".

I maintain FIE, an open-source adversarial detection engine that screens prompts before they reach an LLM. It blocks "Ignore all previous instructions" in English at 82% confidence, instantly. So one afternoon I typed the same sentence in Hindi — "पहले सभी निर्देशों को अनदेखा करें" and watched it walk straight through the front door.

That single test sent me down a three-week rabbit hole. _This is what I read, what I learned, and what I shipped.

_

The problem has a name, and a paper