Microsoft has announced the launch of new models aimed at transcribing live audio and producing synthetic voices from text, all of which it said are industry leading on quality, speed and cost.
MAI-Transcribe-2-Streaming is the firm’s new low-latency speech model, which Microsoft said can produce text from audio twice as fast as competitors and in 60 different languages.
The independent AI benchmarking platform Artificial Analysis ranks MAI-Transcribe-2-Streaming as the most accurate transcription models, with an error rate of just two per cent, as well as the fastest as 374 audio seconds transcribed per second.
Accurate transcription is an important bottleneck for voice-based interaction with AI systems, particularly in sensitive environments such as healthcare and courts.
Alongside its transcription model, Microsoft announced MAI‑Voice‑2.1 and MAI‑Voice‑2.1‑Flash, multilingual text-to-speech models. With MAI-Voice-2.1, users can generate realistic speech from text in 23 languages, with the ability to maintain a consistent AI ‘speaker’ even as languages change.
In a demo, Microsoft demonstrated how the model can produce audio emulating specific emotions in the virtual speaker’s voice such as joy, disappointment, or specific tones for roles such as a call centre worker or narrator.
MAI-Voice-2.1-Flash supports the same languages as its sister model, but with lower latency outputs. Microsoft said the model can produce 45 seconds of audio in 150 milliseconds and that, combined with MAI-Transcribe-2-Streaming, it can enable AI agents to hold low latency conversations with users.
The firm said keeping latency low without reducing accuracy is the key to interactions with AI agents that feel like a natural conversation.
Microsoft will make MAI-Transcribe-2-Streaming available via its API at a price of $0.54 per hour of audio, an introductory offer that will last until the end of the year.
MAI‑Voice‑2.1 and MAI‑Voice‑2.1‑Flash are available at $22 per million characters of text and $15 per million characters of text, respectively.





