Descript has shortened the time needed to test a newly released AI model from as much as a week to one or two hours. The audio and video editing company now runs candidates through its own evaluation harness without first asking an engineer to build a separate provider integration.
Previously, Descript maintained direct connections to OpenAI, Anthropic and Google, along with its own fallback logic. Adding another model required only a few hours of engineering, but scheduling that work could take days. That queue limited which models the team considered for Underlord, its video-editing agent.
The new process uses OpenRouter as one interface to multiple providers. Product staff can request an evaluation through Slack, where an Anthropic integration starts the tests and opens code changes for human review. Descript says it now evaluates models several times a week and can move a successful candidate into production within hours. Most tested models still do not ship.
The company has not handed all hosting decisions to a shared endpoint. It can bring its own provider key for a dedicated deployment and configure other hosts as fallbacks while keeping the selected model constant. The case study comes from OpenRouter, so it reflects a customer and vendor account rather than an independent benchmark, but it illustrates how standardized access can turn model selection into a repeatable product workflow.