A new arXiv paper proposes a benchmark for measuring trust between AI agents. The authors use a cooperative survival game where checking a teammate’s work costs resources, while trusting a wrong answer can be fatal.

The design turns verification behavior into an observable signal. If agents verify less after working with a reliable teammate, that can indicate trust formation; renewed checking can show trust breakage or recovery.

The work is relevant as multi-agent systems become more common. Governance, safety, and coordination will depend not only on individual model quality, but on how agents decide when to rely on each other.