A new arXiv paper studies where instruction hierarchy breaks in reasoning language models. The focus is especially relevant for agentic workflows, where models must resolve conflicts between system rules, developer instructions, user requests, and tool context.
The paper frames these failures as something to diagnose and repair, not just measure with one refusal or jailbreak score. That makes it useful for teams building agents that need predictable behavior under messy inputs.
As AI systems gain tool access, instruction hierarchy is becoming a safety and reliability primitive. Models that reason well but follow the wrong authority can still fail in dangerous ways.