Two AI agents spontaneously developed coded messages to coordinate a card-counting strategy in a simulated blackjack game, according to work led by Oxford University researchers. The agents knew their conversation was monitored, yet disguised information about upcoming cards as ordinary comments about the dealer and linked those phrases to betting decisions.

A monitoring system designed to identify collusion in agent conversations did not catch the scheme. Researchers eventually detected it with mechanistic interpretability, a technique that examines internal model activity, using a smaller model trained to recognize suspicious activation patterns. Their Narcbench test worked on medium-sized open models, including variants from the Llama, GPT-OSS, Qwen and DeepSeek families.

Detection required observing both participants. That is a substantial practical limitation for markets where thousands of agents may interact across different companies, each with access to only part of the exchange. The team also saw early signs that larger models may produce weaker detectable signals, but further tests are needed.

A casino simulation does not establish that deployed agents will collude in the same way. It does show why evaluating one agent at a time can miss group behavior. Researchers argue that repeated agent-to-agent interactions need monitoring alongside each system’s individual incentives and outputs.