Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Note
This feature is currently in public preview. This preview is provided without a service-level agreement, and is not recommended for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
MAI-Voice-2.1 produces natural, expressive speech from text or a short reference clip, with built-in guardrails ensuring only authorized, consented voices can be used. It delivers stable, high-fidelity output that preserves speaker consistency across audiobooks, podcasts, and lectures in 23 different languages.
The following models are supported:
MAI-Voice-2.1MAI-Voice-2.1-Flash
| Model | Voice Count | Key Characteristics | Best For |
|---|---|---|---|
| MAI-Voice-2.1-Flash | Prebuilt voices across 23 languages | Ultra-fast low-latency, emotionally rich, highly expressive, multilingual, supports 23 languages, instant voice cloning (gated), fine-grained emotion control via SSML | Real-time voice agents and assistants, low-latency call center/IVR flows, multilingual interactive experiences |
| MAI-Voice-2.1 | Prebuilt voices across 23 languages | Emotionally rich, highly expressive, high-fidelity, multilingual, supports 23 languages, instant voice cloning (gated), long-form generation with speaker consistency, fine-grained emotion control via SSML | Expressive long-form content, educational content, audiobooks/podcasts, voice overs |
Model details
MAI‑Voice‑2.1‑Flash is a text‑to‑speech model built for fast, low‑latency generation. It produces high‑fidelity, natural, and expressive speech across 23 languages and supports gated instant voice cloning, all while being optimized for real‑time responsiveness. Its human‑like intonation, rhythm, and emotional nuance make it ideal for voice agents, assistants, and other interactive scenarios where latency and cost are critical.
You can integrate with MAI-Voice-2.1-Flash using the Azure Speech SDK via SSML, and also via Voice Live.
Key features
| Key features | Description |
|---|---|
| Ultra-fast low-latency synthesis | Built for real-time text-to-speech with very low latency, suitable for interactive voice scenarios. |
| High-fidelity natural synthesis | Produces natural, expressive, emotionally rich, and high-clarity voice output with human-like rhythm and intonation. |
| Multilingual support | Supports synthesis across 23 languages. |
| Emotion and style control | Developers can influence speaking style by using SSML with mstts:express-as and style, enabling control over emotions such as joy, excitement, empathy, and more. |
| Voice prompting with instant cloning (gated) | Matches a consented reference voice from a short audio clip (5-60 seconds) without additional training. |
| Voice library | Includes licensed curated voices across the supported languages that work out of the box for rapid deployment. |
| Real-time agent optimization | Optimized for voice agents, assistants, IVR, and call-center interactions where responsiveness is critical. |
SSML example
Use this voice name for MAI-Voice-2.1-Flash:
The same suffix applies to any supported prebuilt voice.
Basic SSML
The following example uses Harper with MAI-Voice-2.1-Flash:
Hello world. This is a sample from MAI Voice.
Multilingual SSML
The following SSML synthesizes a greeting in Spanish (Mexico) by using es-MX-Valeria:MAI-Voice-2.1-Flash.
Hola, esta es una muestra de MAI Voice 2.1 Flash.
Expressive control with SSML mstts:express-as
Use the style attribute to control expression:
Welcome to Microsoft Build. MAI Voice 2.1 Flash supports multilingual expressive synthesis.
Prerequisites
- An Azure subscription. You can create one for free.
- A Microsoft Foundry resource for Speech in the Azure portal.
- The Speech resource key and region. After your Speech resource is deployed, select Go to resource to view and manage keys. For the current list of supported regions, see Speech service regions.
Availability and regions
You can access MAI-Voice-2.1 and MAI-Voice-2.1-Flash globally. Azure serves the models from the following regions, and routes requests to them:
| Region | Region identifier | Availability |
|---|---|---|
| France Central | francecentral |
Available |
| East Asia | eastasia |
Available |
| Southeast Asia | southeastasia |
Available |
| East US | eastus |
Available |
| Canada Central | canadacentral |
Available |
| East US 2 | eastus2 |
Available |
| West US | westus |
Available |
| West Europe | westeurope |
Available |
| North Europe | northeurope |
Available |
| West US 2 | westus2 |
Available |
| West US 3 | westus3 |
Available |
| Central India | centralindia |
Available |
| Sweden Central | swedencentral |
Available |
| Japan East | japaneast |
Available |
Pricing
Find pricing information at https://azure.microsoft.com/pricing/details/speech/.
Usage: Available for third-party developers. Microsoft holds full licensing rights for commercial use.
Choose your preferred usage method
MAI-Voice models use the same Azure Speech API and SDK as other Azure neural and HD voices. Choose a portal, API, or SDK tab, and use a MAI-Voice name in the SSML voice element. For available names, see Managed voices and styles.
Try a MAI-Voice voice in the Foundry portal:
- Go to the Text to speech feature page and select Open in playground.
- Select a
MAI-Voice-2.1-Flashvoice from the voice dropdown. - Enter sample text in the text box.
- Select Play to hear the synthesized speech.
Send an SSML POST request to the cognitiveservices/v1 endpoint of your Speech resource.
Set these environment variables:
export SPEECH_KEY=""
export SPEECH_REGION=""
Run the following command:
curl -X POST \
"https://${SPEECH_REGION}.tts.speech.microsoft.com/cognitiveservices/v1" \
--header "Content-Type: application/ssml+xml" \
--header "X-Microsoft-OutputFormat: audio-24khz-160kbitrate-mono-mp3" \
--header "Ocp-Apim-Subscription-Key: ${SPEECH_KEY}" \
--data '
Hello, this is a sample from MAI Voice.
' \
--output output.mp3
On success, an output.mp3 file is saved to the current directory.
Install the Speech SDK for Python, and set the SPEECH_KEY and SPEECH_REGION environment variables.
The following example synthesizes SSML to an MP3 file:
import os
import azure.cognitiveservices.speech as speechsdk
speech_config = speechsdk.SpeechConfig(
subscription=os.environ["SPEECH_KEY"],
region=os.environ["SPEECH_REGION"],
)
speech_config.set_speech_synthesis_output_format(
speechsdk.SpeechSynthesisOutputFormat.Audio24Khz160KBitRateMonoMp3
)
audio_config = speechsdk.audio.AudioOutputConfig(filename="output.mp3")
synthesizer = speechsdk.SpeechSynthesizer(
speech_config=speech_config,
audio_config=audio_config,
)
ssml = """
Hello, this is a sample from MAI Voice.
"""
result = synthesizer.speak_ssml_async(ssml).get()
if result.reason != speechsdk.ResultReason.SynthesizingAudioCompleted:
raise RuntimeError(f"Speech synthesis failed: {result.reason}")
On success, an output.mp3 file is saved to the current directory.
Install the Speech SDK for C#, and set the SPEECH_KEY and SPEECH_REGION environment variables.
The following example synthesizes SSML to an MP3 file:
using System;
using System.IO;
using Microsoft.CognitiveServices.Speech;
var speechConfig = SpeechConfig.FromSubscription(
Environment.GetEnvironmentVariable("SPEECH_KEY"),
Environment.GetEnvironmentVariable("SPEECH_REGION")
);
speechConfig.SetSpeechSynthesisOutputFormat(
SpeechSynthesisOutputFormat.Audio24Khz160KBitRateMonoMp3
);
using var synthesizer = new SpeechSynthesizer(speechConfig);
const string ssml = """
Hello, this is a sample from MAI Voice.
""";
using var result = await synthesizer.SpeakSsmlAsync(ssml);
if (result.Reason != ResultReason.SynthesizingAudioCompleted)
{
throw new InvalidOperationException($"Speech synthesis failed: {result.Reason}");
}
await File.WriteAllBytesAsync("output.mp3", result.AudioData);
On success, an output.mp3 file is saved to the current directory.
Install the Speech SDK for JavaScript:
npm install microsoft-cognitiveservices-speech-sdk
Set the SPEECH_KEY and SPEECH_REGION environment variables. The following Node.js example synthesizes SSML to an MP3 file:
const fs = require("fs");
const sdk = require("microsoft-cognitiveservices-speech-sdk");
const speechConfig = sdk.SpeechConfig.fromSubscription(
process.env.SPEECH_KEY,
process.env.SPEECH_REGION
);
speechConfig.speechSynthesisOutputFormat =
sdk.SpeechSynthesisOutputFormat.Audio24Khz160KBitRateMonoMp3;
const synthesizer = new sdk.SpeechSynthesizer(speechConfig);
const ssml = `
Hello, this is a sample from MAI Voice.
`;
synthesizer.speakSsmlAsync(
ssml,
(result) => {
fs.writeFileSync("output.mp3", Buffer.from(result.audioData));
synthesizer.close();
},
(error) => {
synthesizer.close();
throw error;
}
);
On success, an output.mp3 file is saved to the current directory.
Install the Speech SDK for Java, and set the SPEECH_KEY and SPEECH_REGION environment variables.
The following example synthesizes SSML to an MP3 file:
import com.microsoft.cognitiveservices.speech.ResultReason;
import com.microsoft.cognitiveservices.speech.SpeechConfig;
import com.microsoft.cognitiveservices.speech.SpeechSynthesisOutputFormat;
import com.microsoft.cognitiveservices.speech.SpeechSynthesisResult;
import com.microsoft.cognitiveservices.speech.SpeechSynthesizer;
import java.nio.file.Files;
import java.nio.file.Path;
public class MaiVoiceSynthesis {
public static void main(String[] args) throws Exception {
SpeechConfig speechConfig = SpeechConfig.fromSubscription(
System.getenv("SPEECH_KEY"),
System.getenv("SPEECH_REGION")
);
speechConfig.setSpeechSynthesisOutputFormat(
SpeechSynthesisOutputFormat.Audio24Khz160KBitRateMonoMp3
);
String ssml = """
Hello, this is a sample from MAI Voice.
""";
try (SpeechSynthesizer synthesizer = new SpeechSynthesizer(speechConfig);
SpeechSynthesisResult result = synthesizer.SpeakSsmlAsync(ssml).get()) {
if (result.getReason() != ResultReason.SynthesizingAudioCompleted) {
throw new IllegalStateException(
"Speech synthesis failed: " + result.getReason()
);
}
Files.write(Path.of("output.mp3"), result.getAudioData());
} finally {
speechConfig.close();
}
}
}
On success, an output.mp3 file is saved to the current directory.
Custom voice - Personal Voice/Instant Voice Cloning (gated access)
Developers can create a custom voice in Microsoft Foundry across all supported languages by using just a short reference clip, with no retraining or fine-tuning required. By using only a few seconds of audio (recommended: 5-60 seconds), MAI-Voice models generate high-quality speech that matches the speaker's identity, making it easy for companies to bring their own brand voice into products without maintaining a separate voice model.
All MAI-Voice models support Instant Voice Cloning. Only authorized, licensed voices can be synthesized in production. No unlicensed voice cloning is possible. To gain access to this feature:
- Apply for gated access through Azure AI Custom Neural Voice and Custom Avatar Limited Access Review.
- Once approved, access personal voice APIs at cognitive-services-speech-sdk/samples/custom-voice.
- Upload audio consent and prompt to create a personal voice.
- Synthesize text by using the created voice and a MAI-Voice model. Select a model tab for the corresponding SSML.
The following SSML example uses MAI-Voice-2.1-Flash:
I'm happy to hear that you find me amazing and that I have made your trip planning easier and more fun.
Prebuilt voices
Managed voices and styles
All managed voices in the following table support both MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Use the complete voice ID with the selected model suffix in SSML. For example, use en-US-Harper:MAI-Voice-2.1 or en-US-Harper:MAI-Voice-2.1-Flash.
| Voice ID | Locale | Language | Gender | Supported models | Supported styles |
|---|---|---|---|---|---|
cs-CZ-Grant |
cs-CZ |
Czech | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
cs-CZ-Harper |
cs-CZ |
Czech | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
da-DK-Grant |
da-DK |
Danish | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, customer_call_center, educational, narrator, neutral |
da-DK-Harper |
da-DK |
Danish | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
de-DE-Grant |
de-DE |
German | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
de-DE-Harper |
de-DE |
German | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
de-DE-Klaus |
de-DE |
German (Germany) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
de-DE-Mia |
de-DE |
German (Germany) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
en-AU-Isla |
en-AU |
English (Australia) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
en-GB-Emily |
en-GB |
English (United Kingdom) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, angry, audiobook, confused, customer_call_center, disgusted, educational, embarrassed, excited, fearful, happy, jealous, joyful, narrator, neutral, sad, surprised |
en-GB-Harry |
en-GB |
English (United Kingdom) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, angry, audiobook, customer_call_center, disgusted, educational, fearful, joyful, narrator, neutral, sad, surprised |
en-IN-Dhruv |
en-IN |
English (India) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
en-IN-Priya |
en-IN |
English (India) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
en-US-Ethan |
en-US |
English (United States) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
en-US-Grant |
en-US |
English (United States) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
en-US-Harper |
en-US |
English (United States) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, angry, audiobook, confused, customer_call_center, determined, educational, embarrassed, excited, happy, hopeful, joyful, narrator, neutral, regretful, relieved, sad, shouting, softvoice, whispering |
en-US-Iris |
en-US |
English (United States) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
en-US-Jasper |
en-US |
English (United States) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
en-US-Olivia |
en-US |
English (United States) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
en-US-Sage |
en-US |
English (United States) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
es-ES-Marta |
es-ES |
Spanish (Spain) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, neutral, nostalgic, reflective, saddisappointed, serious |
es-MX-Alejo |
es-MX |
Spanish (Mexico) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
es-MX-Grant |
es-MX |
Spanish (Mexico) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
es-MX-Harper |
es-MX |
Spanish (Mexico) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
es-MX-Valeria |
es-MX |
Spanish (Mexico) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
fi-FI-Grant |
fi-FI |
Finnish | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
fi-FI-Harper |
fi-FI |
Finnish | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
fr-FR-Grant |
fr-FR |
French | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
fr-FR-Harper |
fr-FR |
French | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
fr-FR-Marc |
fr-FR |
French (France) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
fr-FR-Soleil |
fr-FR |
French (France) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
hi-IN-Arjun |
hi-IN |
Hindi (India) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, sad, surprised |
hi-IN-Dhruv |
hi-IN |
Hindi (India) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
hi-IN-Grant |
hi-IN |
Hindi | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
audiobook, neutral |
hi-IN-Harper |
hi-IN |
Hindi | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, customer_call_center, educational, narrator, neutral |
hi-IN-Kavya |
hi-IN |
Hindi (India) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
hi-IN-Priya |
hi-IN |
Hindi (India) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
hu-HU-Bence |
hu-HU |
Hungarian (Hungary) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
hu-HU-Grant |
hu-HU |
Hungarian | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
hu-HU-Harper |
hu-HU |
Hungarian | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
hu-HU-Levente |
hu-HU |
Hungarian (Hungary) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
hu-HU-Lilla |
hu-HU |
Hungarian (Hungary) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
hu-HU-Reka |
hu-HU |
Hungarian (Hungary) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
id-ID-Grant |
id-ID |
Indonesian | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
id-ID-Harper |
id-ID |
Indonesian | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
it-IT-Grant |
it-IT |
Italian | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
audiobook, educational, neutral |
it-IT-Harper |
it-IT |
Italian | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, customer_call_center, narrator, neutral |
it-IT-Luca |
it-IT |
Italian (Italy) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
it-IT-Rosa |
it-IT |
Italian (Italy) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
ko-KR-Grant |
ko-KR |
Korean | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
ko-KR-Haena |
ko-KR |
Korean (Korea) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, neutral, regretful, relieved, sad, softvoice, surprised |
ko-KR-Harper |
ko-KR |
Korean | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
ko-KR-Junho |
ko-KR |
Korean (Korea) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, neutral, relieved, sad, softvoice |
nb-NO-Grant |
nb-NO |
Norwegian (Bokmål) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
educational, neutral |
nb-NO-Harper |
nb-NO |
Norwegian (Bokmål) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, narrator, neutral |
nl-NL-Grant |
nl-NL |
Dutch | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, customer_call_center, educational, narrator, neutral |
nl-NL-Harper |
nl-NL |
Dutch | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
nl-NL-Sander |
nl-NL |
Dutch (Netherlands) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, neutral, nostalgic, reflective, saddisappointed, serious |
pl-PL-Grant |
pl-PL |
Polish | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
pl-PL-Harper |
pl-PL |
Polish | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
pt-BR-Caio |
pt-BR |
Portuguese (Brazil) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
pt-BR-Grant |
pt-BR |
Portuguese (Brazil) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, customer_call_center, educational, neutral |
pt-BR-Harper |
pt-BR |
Portuguese (Brazil) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
audiobook, narrator, neutral |
pt-BR-Luana |
pt-BR |
Portuguese (Brazil) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
pt-BR-Pedro |
pt-BR |
Portuguese (Brazil) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
confused, determined, embarrassed, excited, happy, hopeful, joyful, neutral, regretful, relieved, sad, softvoice, surprised |
pt-BR-Rafael |
pt-BR |
Portuguese (Brazil) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, neutral, regretful, relieved, sad, softvoice, surprised |
pt-PT-Grant |
pt-PT |
Portuguese (Portugal) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
pt-PT-Harper |
pt-PT |
Portuguese (Portugal) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
pt-PT-Rui |
pt-PT |
Portuguese (Portugal) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, neutral, regretful, relieved, sad, softvoice, surprised |
ro-RO-Andrei |
ro-RO |
Romanian (Romania) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
ro-RO-Elena |
ro-RO |
Romanian (Romania) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
ro-RO-Grant |
ro-RO |
Romanian | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, customer_call_center, educational, narrator, neutral |
ro-RO-Harper |
ro-RO |
Romanian | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
audiobook, neutral |
ro-RO-Ioana |
ro-RO |
Romanian (Romania) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
ro-RO-Radu |
ro-RO |
Romanian (Romania) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
ru-RU-Grant |
ru-RU |
Russian | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, customer_call_center, neutral |
ru-RU-Harper |
ru-RU |
Russian | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
audiobook, educational, narrator, neutral |
ru-RU-Lev |
ru-RU |
Russian (Russia) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, neutral, nostalgic, reflective, saddisappointed, serious |
ru-RU-Masha |
ru-RU |
Russian (Russia) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, neutral, nostalgic, reflective, saddisappointed, serious |
sv-SE-Grant |
sv-SE |
Swedish | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
neutral |
sv-SE-Harper |
sv-SE |
Swedish | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
th-TH-Grant |
th-TH |
Thai | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
th-TH-Harper |
th-TH |
Thai | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
th-TH-Krit |
th-TH |
Thai | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, neutral, nostalgic, reflective, saddisappointed, serious |
th-TH-Nattapong |
th-TH |
Thai | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, neutral, nostalgic, reflective, saddisappointed, serious |
tr-TR-Aydin |
tr-TR |
Turkish (Türkiye) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, neutral, nostalgic, reflective, saddisappointed, serious |
tr-TR-Elif |
tr-TR |
Turkish (Türkiye) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, neutral, nostalgic, reflective, saddisappointed, serious |
tr-TR-Grant |
tr-TR |
Turkish | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
educational, narrator, neutral |
tr-TR-Harper |
tr-TR |
Turkish | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, neutral |
vi-VN-Grant |
vi-VN |
Vietnamese | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, customer_call_center, educational, narrator, neutral |
vi-VN-Harper |
vi-VN |
Vietnamese | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
audiobook, neutral |
zh-CN-Bo |
zh-CN |
Chinese (Simplified) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
zh-CN-Grant |
zh-CN |
Chinese (Simplified) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
zh-CN-Harper |
zh-CN |
Chinese (Simplified) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
agent, audiobook, customer_call_center, educational, narrator, neutral |
zh-CN-Lan |
zh-CN |
Chinese (Simplified) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, disgusted, embarrassed, excited, fearful, happy, joyful, neutral, sad, surprised |
zh-CN-Mei |
zh-CN |
Chinese (Simplified) | Female | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, relieved, sad, shouting, softvoice, surprised, whispering |
zh-CN-Wei |
zh-CN |
Chinese (Simplified) | Male | MAI-Voice-2.1, MAI-Voice-2.1-Flash |
angry, confused, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, neutral, regretful, sad, surprised |
Note
Microsoft adds more locales and managed voices as they become available.
Related content
- For more information about using LLM Speech API, see LLM Speech API.
- MAI-Transcribe in Azure Speech.
- How to customize Voice Live input and output.