OpenRouter has launched an asynchronous API for AI jobs that can wait, trading flexible completion times for lower prices. The Batch API generally charges half the normal per-token rate, although discounts vary by model, and gives providers up to 24 hours to finish a workload.
The service supports more than 70 models and accepts chat, response, message and embedding requests. Suitable jobs include labeling datasets, scoring evaluations, backfilling embeddings and processing ticket backlogs. Each row returns independently, so one bad request does not invalidate the whole batch.
OpenRouter says the two-week beta covered more than 230,000 completed batches. The median finished in seven minutes, 90 percent completed within an hour and 99 percent within 10.3 hours. Submission time affected latency more than batch size, with jobs sent from early morning to noon Pacific taking longer.
There are practical limits. A batch runs through one provider, inputs such as images must use public URLs, and audio, video and OpenRouter’s web-search plugin are not supported. Inputs and results are retained for 30 days unless the customer deletes the batch sooner.