Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Text-to-Speech

Best Text-to-Speech Models

Model rankings updated September 2026 based on real usage data.

Text-to-speech models convert written text into spoken audio for assistants, narration, accessibility tools, voiceovers, and customer support. This collection ranks TTS models by their usage on OpenRouter over the past week. The current top models are Gemini 3.1 Flash TTS Preview, S2.1 Pro Free (free), and Kokoro 82M. Compare voices, latency, pricing, and provider capabilities to find the best speech model for your product.

Browse All ModelsCompare Models

Top Text-to-Speech Models on OpenRouter

Favicon for google

Google: Gemini 3.1 Flash TTS Preview

316M tokens

Gemini 3.1 Flash TTS Preview is a text-to-speech model from Google, and a substantial generational step up from Gemini 2.5 Flash TTS. It takes text input and produces audio output across 70+ languages — nearly 3× the language coverage of its predecessor.

The headline addition is a system of 200+ inline audio tags (e.g. [whispers], [laughs], [excited]) that let developers steer delivery, emotion, and pacing mid-sentence, alongside a "director's chair" workflow in Google AI Studio for defining per-character Audio Profiles and scene-level context. It supports up to two speakers with independent voice and style configuration per speaker, outputs PCM audio at 24 kHz / 16-bit mono, and automatically watermarks all output with SynthID. Context window is 32k tokens.

by google33K context$1/M input tokens$20/M output tokens
Favicon for fish-audio

Fish Audio: S2.1 Pro Free (free)

135M tokens

S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability guarantees.

by fish-audio$0/M input tokens$0/M output tokens
Favicon for hexgrad

hexgrad: Kokoro 82M

87.2M tokens

Kokoro 82M is a lightweight, open-weight text-to-speech model from hexgrad. It converts text to speech across 8 languages (American and British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, and Chinese) using 54 preset voices organized by language and gender. At 82M parameters, it is well-suited for multilingual TTS deployments where footprint and cost efficiency matter.

by hexgrad4K context$0.62/M characters
Favicon for x-ai

SpaceXAI: Grok Voice TTS 1.0

8.32M tokens

Grok Voice TTS 1.0 is a text-to-speech model from SpaceXAI. It converts text into spoken audio across 20+ languages with automatic language detection, and offers five built-in voices (Eve, Ara, Rex, Sal, Leo) covering a range of tones. Inline speech tags allow control over pauses, emphasis, pitch, speed, and vocal style. Output is available in MP3, WAV, PCM, μ-law, and A-law formats at sample rates from 8 kHz to 48 kHz, with up to 15,000 characters per request.

by x-ai15K context$15/M characters
Favicon for deepgram

Deepgram: Flux TTS (free)

3.93M tokens

Flux TTS is a text-to-speech model from Deepgram. It is suited for natural, expressive English speech synthesis across Deepgram's Flux voice catalog.

by deepgram$0/M input tokens$0/M output tokens
Favicon for fish-audio

Fish Audio: S2.1 Pro

3.07M tokens

S2.1 Pro is a production-oriented text-to-speech model from Fish Audio. It is suited for multilingual voice applications, expressive narration, and dialogue synthesis, with open-ended natural-language controls for speaking style and emotion.

by fish-audio$15/M UTF-8 bytes
Favicon for microsoft

Microsoft: MAI-Voice-2

2.74M tokens

MAI-Voice-2 is an expressive text-to-speech model from Microsoft. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales, fine-grained control of tone and delivery, multi-speaker generation, and voice prompting from short audio clips without fine-tuning. The model prioritizes naturalness and expressivity over latency-critical generation.

by microsoft$22/M characters
Favicon for qwen

Qwen: Qwen-Audio-3.0-TTS Flash

1.2M tokens

Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

by qwen$15/M characters
Favicon for deepgram

Deepgram: Aura-2

964K tokens

Aura-2 is a multilingual text-to-speech model from Deepgram. It supports Deepgram’s canonical Aura-2 voice catalog for speech synthesis across multiple languages.

by deepgram$30/M characters
Favicon for microsoft

Microsoft: MAI-Voice-2-Flash

832K tokens

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15 languages and 18 locales, with fine-grained control over tone and delivery. Voice prompting and cloning require Microsoft-approved access and appropriate speaker consent.

by microsoft$15/M characters
Favicon for qwen

Qwen: Qwen-Audio-3.0-TTS Plus

597K tokens

Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

by qwen$20/M characters
Favicon for mistralai

Mistral: Voxtral Mini TTS

468K tokens

Voxtral Mini TTS is Mistral's text-to-speech model featuring zero-shot voice cloning and multilingual support. It converts text input into natural-sounding audio output.

by mistralai4K context$16/M characters
Favicon for minimax

MiniMax: Speech 2.8 Turbo

347K tokens

MiniMax Speech 2.8 Turbo is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.

by minimax$60/M characters
Favicon for minimax

MiniMax: Speech 2.8 HD

265K tokens

MiniMax Speech 2.8 HD is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.

by minimax$100/M characters
Favicon for fish-audio

Fish Audio: S2 Pro

82K tokens

S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.

by fish-audio$15/M UTF-8 bytes
Favicon for canopylabs

Canopy Labs: Orpheus 3B

68K tokens

Orpheus 3B is an English text-to-speech model from Canopy Labs, fine-tuned for natural prosody and expressive delivery. It offers 7 preset voices and is suited for narration, voice assistants, and interactive applications where naturalistic speech is a priority.

by canopylabs4K context$7/M characters
Favicon for fish-audio

Fish Audio: S1

32K tokens

S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported languages.

by fish-audio$15/M UTF-8 bytes
Favicon for sesame

Sesame: CSM 1B

11K tokens

CSM 1B is a conversational speech model from Sesame. It accepts text input and produces English speech output, with voice options spanning conversational and read-speech styles. At 1B parameters, it is suited for dialogue-oriented applications such as voice assistants and interactive agents.

by sesame4K context$7/M characters
Favicon for zyphra

Zyphra: Zonos v0.1 Transformer

Zonos v0.1 Transformer is a text-to-speech model from Zyphra built on a pure transformer architecture. It offers the same American and British English voice coverage as the Hybrid variant, and is suited for deployments where a transformer-only inference stack is preferred.

by zyphra4K context
Favicon for zyphra

Zyphra: Zonos v0.1 Hybrid

Zonos v0.1 Hybrid is a text-to-speech model from Zyphra built on a hybrid architecture. It produces English speech output with coverage across American and British accents in male and female voices. It is suited for English-language voice applications requiring accent and gender variety.

by zyphra4K context

Explore more collections

  • Free Models
  • Discounted Models
  • Coding
  • Roleplay
  • Vision Models
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Speech-to-Text
  • Embedding Models
  • Rerank Models
  • Distillable Models
  • All collections