Public benchmarks cannot tell a team whether a model change still works for its own users. OpenRouter recommends building a “golden” evaluation set from production inputs, reviewed expected behavior and rubrics, then versioning it alongside prompts and code.
The process starts with a week or two of traffic and metadata such as intent, feature, model and prompt version. Personal information should be scrubbed before examples enter an evaluation harness. Teams then deduplicate repeated requests, cluster similar intents and preserve the real distribution while adding enough rare, high-cost failures to make the set useful.
A domain expert should define the expected output or binary criteria for every case. Two reviewers can annotate a subset independently; disagreement often exposes an unclear rubric. Running the current production model once helps separate genuine failures from mislabeled or inherently ambiguous examples before the set becomes a deployment gate.
OpenRouter suggests about 10 cases for exploring one issue and roughly 100 to 1,000 for a broad regression set, with a smaller subset in pull requests and full runs on release branches or overnight. Synthetic examples can fill known gaps but should remain labeled and secondary. Dataset, rubric and baseline must change together in version control, otherwise a score movement cannot be traced to the model, prompt or test itself.