Editions
Base (0.6B and 1.7B), Presets (0.6B), Voice Design (1.7B)
Languages
10, including Chinese, English, Japanese, Korean, German, French and Spanish
Voices
9 presets, cloned voices or designed voices
Voice design
Yes, with the Voice Design edition
Voice training
Yes, on the 0.6B Base edition
Playback
Plays while it speaks
Made by
Alibaba Qwen

Qwen3-TTS is a family of editions that each do one job well. Avocado recognises which one you have and shows only what it can do.

Presets

The Presets edition has nine ready voices, including two that speak Beijing and Sichuan dialects.

Clone

The Base editions speak with a voice you clone. Give the clip’s words too and you get a closer clone; Avocado saves it as a ready voice almost instantly. Only clone voices you have permission to use.

Design

With the Voice Design edition, a description becomes the voice. Save it and Avocado renders it once, so the same voice also works on the faster Base editions.

Train

The 0.6B Base edition can be taught a voice from your own recordings. Avocado keeps the best result it finds and stops early when more steps won’t help.

Licence

Qwen3-TTS’s licence is set by Alibaba Qwen. Read it on the Qwen3-TTS model page before you publish anything you make.

Questions

Which Qwen3-TTS edition should I get?

Base to clone and train, Presets for nine ready voices, and Voice Design to create voices from a description. A voice designed and saved on Voice Design can also be used on the faster Base editions.

Can I choose the language?

Yes. Each take has a language menu, or let Qwen3-TTS detect it from the text.

Avocado is almost here

Free for Apple silicon Macs running macOS 15 or later. Coming soon.

Coming soon for Mac