OpenRouter has published a continuous-integration pattern that prevents prompt or agent changes from merging when a fixed evaluation set regresses. Test cases live in the repository beside prompts and tool schemas, and a script exits with a failure code when too few model responses satisfy the expected rules.
The guide recommends starting with real production failures and recording both facts an answer must include and statements it must avoid. Each case runs several times, with a majority vote reducing the effect of nondeterministic output. Teams should measure repeated runs on an unchanged branch to establish the normal noise floor before choosing a pass threshold.
A GitHub Actions workflow first checks whether a pull request changed relevant prompts, agent code, tool schemas, tests or evaluation scripts. It then runs the evaluation only when needed. Both the path-checking job and evaluation job should be required status checks; otherwise a failed dependency can skip the evaluation and still appear acceptable.
The approach has limits. Simple string rules miss valid paraphrases and subtle factual errors, while model-based graders add their own cost and uncertainty. Live provider changes can also resemble a code regression, so the example pins requests to one provider. The gate is most useful as a reviewed, versioned safety net rather than proof that an agent is generally reliable.