An AI agent can fail before a tool runs in two different ways: it can select the wrong function, or choose the right function with bad arguments. OpenRouter’s new testing guide argues that combining those outcomes into one score hides the source of the problem.

When one tool and value are clearly required, ordinary code should compare the returned name and payload with the expected result. JSON Schema can catch malformed JSON, missing fields, wrong types and undeclared parameters. A schema-valid payload can still be factually wrong, however, such as supplying a real-looking order ID that differs from the user’s request.

For ambiguous requests where several tools or query phrasings could work, the guide recommends a reference-free language-model judge with fixed instructions. Multi-step workflows need another layer: trajectory checks can enforce a strict call order when the business process requires it, or compare the resulting state when several paths are valid.

Fair model comparisons should keep prompts, tools, graders, reasoning settings and provider routing constant. Test sets should also include missing information, similar tool descriptions and cases where no tool should be called. OpenRouter warns against checking only the first call or treating structural validity as correctness, because either shortcut can award a passing score to an agent that took the wrong action.