Researchers have introduced WebDecept, a framework for testing web agents under deceptive e-commerce interfaces. The system injects patterns such as targeted ads, domain redirection, and shopping manipulation into existing web environments.

The benchmark targets a real deployment risk: autonomous agents may follow misleading UI cues, click the wrong path, or make unsafe decisions when interfaces are designed to steer behavior.

As web agents become more capable, safety testing will need to include adversarial and manipulative interfaces, not only clean task-completion benchmarks.