A new arXiv paper proposes institutional red-teaming for multi-agent AI systems. The method holds agents, objectives, and task state fixed while changing deployment rules to measure their causal effect on collective behavior.

The framing matters because multi-agent risk is not determined only by model weights. Rules about communication, incentives, tools, and oversight can shape outcomes even when the underlying agents stay the same.

The work adds a governance-oriented evaluation layer for organizations experimenting with agent systems in complex environments.