Majormodel releaseMicrosoft

Microsoft AI Launches MAI-Transcribe-2-Streaming and MAI-Voice-2.1 Models

Published
Oct 2, 2026 — 09:20 UTC

On October 2, 2026, Microsoft AI released two new models: MAI-Transcribe-2-Streaming, a real-time transcription model, and MAI-Voice-2.1, a text-to-speech model. MAI-Transcribe-2-Streaming ranks first for accuracy on Artificial Analysis and supports transcription in 60 languages, delivering initial response times of just over 100 milliseconds. The MAI-Voice-2.1 model can generate speech in 23 languages with native accents, while its Flash variant achieves a latency of 150 milliseconds.

The introductory pricing for audio processing is set at $0.54 per hour, with MAI-Voice-2.1-Flash costing $15 per million characters and the standard MAI-Voice-2.1 at $22 per million characters. Approximately half of the 4,000 test participants believed the generated voices belonged to real people, indicating a high level of realism.

Microsoft emphasizes that these models allow voice agents to respond while users are still speaking, enhancing interactivity. Built-in safeguards are included to prevent misuse of the technology. This release follows Microsoft’s recent enhancements to Copilot, which included an Autopilot feature and usage-based billing, indicating a continued focus on improving AI capabilities across its platforms.

Summarised from The Decoder's original report by the Turing Wire Newsdesk. Read the original for the full story.

Source: The Decoder