A new arXiv paper introduces RIFT-Bench, a methodology for red-teaming agentic AI systems whose risks extend beyond traditional prompt attacks. The benchmark represents systems as graphs, then uses that structure to compare different agent architectures.

RIFT-Bench runs in two phases: discovery, which extracts the system structure, and scanning, which deploys adaptive adversarial attacks before producing an evaluation report. The aim is to make security testing less tied to a single implementation or domain.

As agents gain tools, memory, and autonomy, this kind of dynamic evaluation could become more important than static jailbreak tests.