You've just published a post, a video, or a podcast episode you're proud of. Readers are sharing it.
Then someone asks, "Is there a Hindi version?" or "Can I listen to this in French?"
For most creators, that's where it stalls.
Translating the words is the easy part. The voice is harder, since hiring voice artists costs money and re-recording everything yourself takes time you don't have.
That's the gap Microsoft is aiming at with three new models it released on October 1. Let's look at what they do and where they fit.
The three models, in plain terms
Microsoft AI released a set of voice tools, mostly aimed at developers for now:
MAI-Voice-2.1: turns text into expressive, natural speech
MAI-Voice-2.1-Flash: a faster, cheaper version built for live, high-volume use
MAI-Transcribe-2-Streaming: turns speech into text as it's being spoken
The first two handle the speaking. The third handles the listening.
One voice, many languages
The headline feature is that a single voice keeps its identity across languages. Microsoft says the model adapts its pronunciation and delivery to each language, so it doesn't just carry one accent into all of them.
It covers 23 languages. Hindi and English (India) are both on the list, along with French, German, Spanish, Korean, Chinese, Thai, and more.
Other Indian languages like Tamil, Bengali, or Marathi aren't covered yet. If your audience speaks those, this isn't a fit for now.
Cloning your own voice
You can also clone a voice from a short clip of 5 to 60 seconds. Then you could, in theory, have your own voice narrate in other languages.
Access to cloning is gated, though. You need Microsoft's approval, and there are safeguards to prevent misuse.
Where transcription comes in
Voice is only half the picture. Captions matter just as much, especially for video.
MAI-Transcribe-2-Streaming transcribes continuously across 60 languages and detects the language automatically. Microsoft says text starts appearing within a fraction of a second, so it suits live captions as well as recordings.
What it costs
Pricing is simple. It's charged per character, not per minute:
MAI-Voice-2.1: $22 per 1 million characters
MAI-Voice-2.1-Flash: $15 per 1 million characters
MAI-Transcribe-2-Streaming: $0.54 per hour of audio, an introductory price through the end of the year
For a sense of scale, a 1,500-word script is about 9,000 characters. By my estimate, that's roughly $0.20 on the standard model.
What creators could do with it
Put together, the tools open up a few possibilities:
Narrate blog posts or newsletters as audio
Make voiceovers for videos in several languages
Produce audiobook-style content
Add live captions to streams and webinars
The catches
This is early, and it's worth being clear about what it isn't:
It's developer-first. Access is through Azure Speech and Foundry, plus platforms like OpenRouter. I couldn't confirm a simple no-code way for beginners to try it.
It's in preview. Details and pricing may change.
It isn't a dubbing tool. It turns text into speech. You'd still need to handle translation and video editing separately.
Many of the claims are Microsoft's own. Its speed, cost, and accuracy comparisons haven't been independently tested here.
What to watch next
The next things to watch are more languages and easier access. If either arrives, tools like this could become practical for solo creators, not just developers.
For now, it's a good one to understand early, especially if you're thinking about reaching audiences beyond your first language.
Helpful links
100+ Claude Code hacks to ship code 10X faster
Top engineers at Anthropic and OpenAI say AI now writes 100% of their code.
If you're not using AI, you're spending 40 hours doing what they do in 4.
These 100+ Claude Code hacks fix that and help you ship 10x faster.
Sign up for The Code and get:
100+ Claude Code hacks used by top engineers — free
The Code newsletter — learn the latest AI tools, tips, and skills to code faster with AI in 5 minutes a day



