Indirect prompt injection should be evaluated as an adaptive search problem rather than a fixed property of a victim model, according to a new security preprint. These attacks hide instructions in documents, websites or other content that a tool-using agent reads while completing a legitimate task.
The researchers built an attacker with a dedicated harness for inspecting the environment, organizing possible strategies and using feedback from the victim agent to refine later attempts. Across varied tasks, giving that attacker more test-time computing improved both its discovery of vulnerable paths and its ability to exploit them.
Explicit strategy management mattered at larger budgets. Without it, the attacker repeated similar ideas and gains flattened sooner. That result means two evaluations can report very different success rates against the same system simply because one attacker searched longer or managed its attempts more effectively.
The work is an arXiv preprint, and its task set does not represent every browsing or enterprise agent. Its reporting lesson is immediately useful: security claims should include the attacker’s procedure, available tools, number of attempts and compute budget. Defenders should also test adaptive multi-step attacks, because a low failure rate under one-shot prompts may not survive systematic exploration of the full application environment.