AI

Microsoft’s 10-person team built a speech model that undercuts OpenAI, Google and ElevenLabs on every benchmark — at $0.10 an hour

Adrian Kessler

The model is called MAI-Transcribe-2. Microsoft AI, the division led by Mustafa Suleyman, shipped it on September 3 at $0.10 per audio hour — limited through the end of 2026. The previous version, MAI-Transcribe-1.5, cost $0.36 per hour. That 72% price reduction is not a marketing rounding. A contact center processing 100,000 hours of audio per year was paying $36,000 for the prior model; the same workload now costs $10,000.

The benchmark performance is what makes the price unusual. On the FLEURS multilingual evaluation, MAI-Transcribe-2 ranks first across 60 languages with a 5.2% average word error rate — better than OpenAI‘s GPT-Transcribe, Google‘s Gemini 3.5 Transcribe, ElevenLabs Scribe v2, and Meta’s Whisper V3-Large. On the Artificial Analysis leaderboard, which tracks both accuracy and latency together, it sits at second overall and defines what the researchers call the Pareto frontier — the boundary beyond which you cannot improve one metric without degrading the other. The model was trained to sit at the efficient edge of that frontier, not to chase the top of any single-metric table.

Speed is the second claim worth examining. Microsoft says MAI-Transcribe-2 processes audio 10 times faster than GPT-Transcribe, 7 times faster than ElevenLabs Scribe v2, and 5 times faster than Gemini 3.5 Transcribe. For long-form audio — podcast transcription, legal depositions, clinical notes — that difference determines whether a workflow runs in seconds or minutes. The model also does speaker diarization (identifying who spoke when), word-level timestamps, keyword biasing for domain terminology, and code-switching support for languages that mix mid-sentence, specifically citing Hinglish and Spanglish as tested examples.

The team that built it is notable in proportion to what it built. Microsoft describes the core group as 10 people, operating with minimal internal bureaucracy and supported by larger teams handling data acquisition and vendor management. The resulting GPU efficiency claim is specific: MAI-Transcribe-2 runs at half the GPU cost of competing state-of-the-art models. Whether that holds under enterprise load at scale is not verifiable from the announcement, but the architecture claim is precise enough to test.

The model is available now on Microsoft Foundry, the MAI Playground, and Open Router — meaning access does not require an Azure contract or a Microsoft enterprise relationship. That distribution breadth is deliberate. Microsoft AI’s push into direct-to-developer APIs is part of a broader repositioning following the restructuring of the OpenAI partnership. Amendments in October 2025 and April 2026 eliminated the exclusivity and revenue-sharing arrangements that had defined the relationship since 2019. Microsoft now has a direct incentive to compete at the model layer rather than solely distribute OpenAI’s output.

The $0.10 price is marked as limited through end of 2026. Standard pricing after that date has not been disclosed. Buyers evaluating MAI-Transcribe-2 for production workloads are pricing a product whose cost they do not know for calendar year 2027. That is a risk worth naming clearly, even if the current rate makes the model attractive against every visible alternative.

Tags: , , , , ,

Discussion

There are 0 comments.