Models

Microsoft AI Releases Fast MAI Speech and Voice Models

Microsoft AI has launched three new models designed to power ultra-low-latency voice agents, offering rapid speech-to-text transcription and highly realistic multilingual voice cloning.

The Decoder4 days agoModels
Image: The Decoder

Microsoft AI has introduced a real-time transcription model called MAI-Transcribe-2-Streaming. According to Microsoft, the system currently ranks first for accuracy on the Artificial Analysis benchmark. Capable of processing 60 different languages, the technology delivers its initial partial results in slightly more than 100 milliseconds. This rapid processing speed is designed to enable conversational AI agents to formulate responses even while a user is still speaking. To encourage adoption, Microsoft is offering an introductory rate of $0.54 per hour of audio through the end of the year.

Alongside the transcription tool, the company launched two text-to-speech models. The primary model, MAI-Voice-2.1, can speak 23 languages using a single consistent voice while maintaining a native accent for each language. A faster variant, MAI-Voice-2.1-Flash, reduces latency to a mere 150 milliseconds. This flash version is also more cost-effective, priced at $15 per million characters compared to the standard rate of $22 per million characters.

Both of the new voice models feature the ability to clone a specific voice using only a few seconds of reference audio, backed by integrated safeguards to prevent unauthorized use. In a listening test involving approximately 4,000 participants, about half of the respondents believed the generated voices belonged to an actual human being. Developers can access these models through Microsoft Foundry and the MAI Playground, while the two text-to-speech models are also available on the OpenRouter platform.

This is our own summary of reporting by The Decoder

More in Models

Microsoft AI Releases Fast MAI Speech and Voice Models | Latest News