Microsoft AI has released one streaming transcription model and two text-to-speech models aimed at voice agents that must listen and respond with little delay. The products are available through Microsoft Foundry and the MAI Playground, while the speech generators are also listed on OpenRouter.

MAI-Transcribe-2-Streaming supports 60 languages and produces its first partial transcription in slightly more than 100 milliseconds, according to Microsoft. The company says this can let an agent begin processing a request before a speaker finishes. An introductory price of $0.54 per hour of audio applies through the end of 2026.

For generated speech, MAI-Voice-2.1 is designed to use the same voice across 23 languages while giving each one a native accent. Its Flash variant has a claimed latency of 150 milliseconds and costs $15 per million characters, compared with $22 for the standard version. Both voice models can clone a voice from a few seconds of reference audio and include safeguards intended to limit misuse.

Microsoft also reports that roughly half of 4,000 test participants mistook generated voices for real people. The accuracy ranking, latency figures and safety claims come from Microsoft or cited benchmarks and still require evaluation in specific production conditions.