AWS has detailed a SageMaker AI workflow for fine-tuning search agents with multi-turn reinforcement learning. Instead of scoring each answer in isolation, the system evaluates a complete sequence of searches, tool calls and decisions according to whether the final result succeeds.

That distinction matters because an agent’s later choices depend on information retrieved earlier. Supervised fine-tuning requires expensive examples of ideal trajectories, while single-turn reinforcement learning can miss those dependencies. AWS’s approach generates multi-step rollouts and uses a final reward to improve the policy across the whole interaction.

Developers provide the tool loop, environment and reward definition. SageMaker manages serverless execution, trajectory collection and model updates, with per-token pricing rather than a dedicated GPU cluster. Supported optimization methods include PPO, CISPO and importance-sampling losses, paired with several group-based advantage estimators. Training jobs can resume across service time limits, and MLflow records actions and rewards for inspection.

AWS positions the method as a way to give a smaller model environment-specific reliability while retaining lower latency and cost than a frontier model. That outcome depends on reward quality and representative training tasks. A flawed verifier can reward shortcuts, and results from one search environment do not establish equal gains for unrelated agents or tools.