Multilingual, directable speech synthesis

Qwen Audio 3.0 TTS for Expressive Multilingual Speech

Learn how Qwen Audio 3.0 TTS handles 16 languages, natural-language voice direction, inline expression tags, long-form generation, and robust voice cloning.

0 characters
Optional

What Is Qwen Audio 3.0 TTS?

Qwen Audio 3.0 TTS is a production-oriented speech synthesis system designed for consistent content, recognizable speaker identity, natural prosody, controllable delivery, multilingual output, and difficult real-world reference audio. It uses a low-frame-rate speech tokenizer to reduce decoding work while retaining speech content and speaker information.

The system supports both broad instructions and local edits. You can describe an overall performance in natural language, then place tags in the script where a laugh, breath, emotion shift, or other non-verbal detail should occur.

16
supported languages
2
model variants
3 min
one-pass long-form synthesis
86
fine-grained inline tags

Facts summarized from the official Qwen Audio 3.0 TTS release.

How Qwen Audio 3.0 TTS Works

The practical workflow combines a script, a selected or authorized voice, an optional performance instruction, and a model variant matched to the job.

1

Prepare text and a voice

Start with the script you want to synthesize. Select a preset voice or provide an authorized reference recording when you need to preserve a specific speaker identity.

2

Describe the delivery

Write a natural-language instruction for emotion, role, pace, accent, or speaking scenario. Add inline tags such as [giggles], [gasp], or [angry] when a specific phrase needs local control.

3

Choose Flash or Plus

Use Flash for interactive experiences where response time matters. Use Plus for finished narration and production work where naturalness and timbre fidelity take priority.

Qwen Audio 3.0 TTS Key Features

The system combines an efficient speech representation, staged training, directable performance control, broad deployment support, and evaluation across real speech-production conditions.

Low-frame-rate speech tokenizer

A supervised 12.5 Hz speech tokenizer lowers the amount of autoregressive decoding needed for synthesis while retaining the content of the speech and the identifying qualities of the speaker.

Listen to Olivia Lin demonstrate the voice quality
Olivia LinEnglish

Click the record to listen

Progressive training pipeline

The training process combines separate language-model and flow-matching pretraining, joint training with high-quality data annealing, language-model reinforcement learning, flow-matching robustness training, and final reinforcement learning. These stages target content consistency, natural prosody, voice fidelity, audio quality, and robustness together.

Listen to Luna Wang demonstrate the voice quality
Luna WangEnglish

Click the record to listen

Production-level controllability

Qwen Audio 3.0 TTS follows free-form instructions for role, emotion, speaking style, rate, timbre, and accent. It also adds 86 fine-grained inline tags for phrase-level and word-level control, including transitions, laughter, breathing, coughing, and sighing.

Listen to Alexander Hu demonstrate the voice quality
Alexander HuEnglish

Click the record to listen

Broad and robust deployment coverage

Qwen Audio 3.0 TTS supports 16 languages, including seven newly added languages. The model also handles difficult text normalization, one-pass synthesis up to about three minutes, and degraded reference prompts without a separate denoising mode. Speaker fine-tuning and vocoder super-resolution support target-voice adaptation and 48 kHz output.

Listen to Nora Hu demonstrate the voice quality
Nora HuEnglish

Click the record to listen

Comprehensive speech evaluation

The reported evaluation covers zero-shot voice cloning, multilingual and cross-lingual synthesis, free-form instruction following, fine-grained control, text normalization, long-form generation, and robustness to difficult reference audio. Results combine objective benchmarks with arena-style human evaluation.

Listen to Adrian Gao demonstrate the voice quality
Adrian GaoEnglish

Click the record to listen

Qwen Audio 3.0 TTS Flash vs Plus

Both variants share the Qwen Audio TTS family, but they prioritize different production needs. Choose based on whether response time or finished-audio quality matters more.

Comparison of Qwen Audio 3.0 TTS Flash and Plus
ComparisonFlashReal-time interactionPlusHigher-quality generation
PositioningTuned for real-time interaction and conversational experiences.Tuned for finished audio where generation quality takes priority.
Best forInteractive speechFinished narration
Primary advantageLower response delayHigher priority on naturalness
Response profileFirst-packet latency at the 300 ms level, as described in the release article.Prioritizes naturalness, speaker similarity, and timbre fidelity over the fastest first packet.
Typical usesVoice assistants, interactive characters, and live experiencesDubbing, narration, and production audio

Variant positioning and latency wording are based on the Tongyi Lab release article.

Listen to Qwen Audio 3.0 TTS

Preview eight speech styles across narration, multilingual delivery, character performance, podcasts, and production voiceover.

Expressive Narration audio scene

Expressive Narration

Audiobook and long-form speech

Multilingual Broadcast audio scene

Multilingual Broadcast

Global news and announcements

Fantasy Character audio scene

Fantasy Character

Expressive character performance

Commercial Voiceover audio scene

Commercial Voiceover

Product and campaign audio

Warm Storytelling audio scene

Warm Storytelling

Stories for families and children

Calm Guidance audio scene

Calm Guidance

Meditation and wellbeing audio

Podcast Conversation audio scene

Podcast Conversation

Natural multi-speaker dialogue

Dramatic Dialogue audio scene

Dramatic Dialogue

Games and cinematic scenes

Qwen Audio 3.0 TTS Use Cases

The same control system can support localized content, long-form narration, and expressive characters without reducing every performance to a fixed preset.

Multilingual dubbing and localization

Keep a consistent speaker identity while generating localized dialogue in multiple languages. Natural-language direction can adapt the performance to a documentary, product video, drama, or short-form clip without building a separate acoustic preset for every scene.

  • Cross-lingual voice consistency
  • Emotion and scene instructions
  • Inline non-verbal cues

Narration and long-form content

Generate audiobook excerpts, training material, explainers, and editorial narration with control over pace and tone. One-pass long-form synthesis reduces the need to split every passage into short fragments before generation.

  • Up to about three minutes in one pass
  • Text normalization support
  • Plus model for finished audio

Games and expressive characters

Direct character voices with descriptions of role, mood, accent, and speaking style. Inline tags help place laughter, breaths, gasps, and emotional changes at the exact point where the script calls for them.

  • Character and role direction
  • Phrase-level expression
  • Repeatable prompt patterns
Qwen3 TTS Pricing

Choose Your Qwen3 TTS Credit Pack

Buy credits once and use them across Qwen text to speech, AI Voice Design, AI Voice Cloning, and multilingual voice generation. Your credits do not expire.

Free

$0forever
2 Credits
About 200 characters
MP3 and WAV export
No credit card required
Personal use only
Standard processing

Basic

$9.9one-time
990 Credits
About 99,000 characters
$0.010 per 100 characters
Full commercial rights
Zero-shot voice cloning
MP3 and WAV export
Email support
Credits never expire
Most Popular

Pro

$29.9one-time
3,700 Credits
About 370,000 characters
$0.008 per 100 characters
Save 50% vs Basic
Long-form continuation
Latest voice model
MP3 and WAV export
Commercial license
Email support
Credits never expire

Business

$49.9one-time
12,400 Credits
About 1,240,000 characters
$0.004 per 100 characters
Save 80% vs Basic
Long-form continuation
Latest voice model
MP3 and WAV export
Commercial license
Email support
Credits never expire

One-time credit packs with no recurring subscription

Pay onceCredits never expireSecure paymentsEmail support: support@qwentts.net

Qwen Audio 3.0 TTS FAQ

Direct answers about languages, model variants, style control, voice cloning, and access.

Qwen Audio 3.0 TTS is a production-oriented text-to-speech system from Qwen Audio, Token Foundry, and Alibaba Group. It combines multilingual speech generation, voice cloning, natural-language direction, inline control tags, and long-form synthesis.

Qwen Audio 3.0 TTS supports 16 languages: Chinese, Arabic, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, and Vietnamese. The announcement notes that full language support is rolling out.

Flash is tuned for interactive use and lower first-packet latency. Plus is tuned for higher-quality generation where naturalness and timbre fidelity are more important than the fastest response.

Yes. Qwen Audio 3.0 TTS supports voice cloning and cross-lingual generation. Only use recordings you own or have permission to use, and disclose synthetic speech where applicable.

You describe how the line should sound in normal language, including emotion, role, pace, timbre, accent, or scenario. The model uses that instruction to guide the overall performance without requiring manual acoustic settings.

Inline tags are markers placed inside the target text to control a local moment. The release adds 86 tags for details such as laughter, breathing, coughing, sighing, emotion changes, and other non-verbal events.

This page provides a Qwen Audio 3.0 TTS generator for English system voices. Choose Flash or Plus, preview a voice, set the delivery instruction and audio controls, then generate and download the result.

Explore Qwen Audio 3.0 TTS

Read the first-party release for technical results and audio demos, or try the Qwen3 TTS tools available on this website for text-to-speech, voice design, and voice cloning workflows.