A new arXiv paper proposes treating AI behavior forecasting as its own learning task.

The authors argue that for large reasoning models, conventional explanations do not easily map onto long trajectories. Their alternative is to train Behavior Forecasters that use a reasoning trajectory to predict future model behavior.

The idea is useful because trust often depends on knowing how a system will act on new inputs, not just reading a plausible explanation after the fact.