ModelInfo
English
All models

MAI-Voice-2.1

Highest-fidelity expressive text-to-speech model across 23 languages, used via Azure Speech SSML.

Microsoftmai-voice-2.1Report incorrect data

Specs

Context
—
Max output
—
Input price
—
Output price
—
Released
—

Context

Not published

Pricing

Not published

Modalities

Input
Text
Output
Audio

API

API types
audio

Reasoning

Not published

Info

Status
Preview
Released
—
Knowledge cutoff
—

Features

Not published
Source · official docsVerified 2026-10-11

How to call

1 provider

microsoftmodel = MAI-Voice-2.1

audio
POSThttps://<region>.tts.speech.microsoft.com/cognitiveservices/v1

JSON

Standard format
mai-voice-2.1.json
{
  "id": "mai-voice-2.1",
  "object": "model",
  "created": null,
  "owned_by": "microsoft",
  "name": "MAI-Voice-2.1",
  "api": {
    "types": [
      "audio"
    ]
  },
  "limits": {
    "context": null,
    "input": null,
    "output": null
  },
  "modalities": {
    "input": [
      "text"
    ],
    "output": [
      "audio"
    ]
  },
  "reasoning": {
    "supported": null,
    "efforts": []
  },
  "pricing": null,
  "features": [],
  "info": {
    "status": "preview",
    "release_date": null,
    "knowledge_cutoff": null,
    "description": "Highest-fidelity expressive text-to-speech model across 23 languages, used via Azure Speech SSML.",
    "docs": "https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices",
    "verified_at": "2026-10-11"
  }
}

Official sources

2