AI agents can regress even when application code does not change. A revised system prompt, a new model behind an alias, an altered tool schema or a retrieval update can change which actions an agent takes while leaving its answer fluent and plausible.

OpenRouter’s testing guide recommends a locked set of representative production cases with a written contract for each one. Structural assertions should specify required tools and arguments. Hard invariants should capture rules that must never break, such as sending a high-value refund to a human instead of approving it automatically.

Exact text comparison is a poor fit because two valid agent answers may look different. Model comparisons should instead hold prompts, tools, cases, judges and inference settings constant while changing only a concrete model version. Using a moving “latest” alias makes results impossible to reproduce; responses should also record the actual model that served them.

The baseline needs scrutiny before the candidate. A case that fails both versions may expose a broken test rather than a regression, while an invariant that fails only on the candidate should block release. Because these evaluations call live models and cost money, the guide suggests a separate measured CI job triggered by changes to prompts, models, tools or retrieval configuration.