Voices
10 preset speakers, plus your own
Languages
English and Chinese
Direction
Written direction, like "warm and unhurried"
Voice cloning
From 1 to 30 seconds of speech with a transcript
Voice training
Yes, from 1 to 30 minutes of speech
Speed
About 0.3 s to first sound on an M3 Max
Made by
BreezeBlue
Breeze TTS 2 speaking a script with the direction field and guidance control below it

Breeze TTS 2 is one of the most capable models Avocado works with: it takes direction, clones, designs and can be taught a new voice.

Speakers and direction

Choose one of its ten preset speakers and, if you like, add a direction for the take: “warm and unhurried”, “whispering”, “excited, a bit faster”. Guidance sets how strongly the direction is followed. Sound tags such as laughs and sighs can go right in the script.

Clone

Clone a voice from 1 to 30 seconds of clear speech with an exact transcript of what was said. Only clone voices you have permission to use; Avocado asks you to confirm before it makes a clone.

Train

Teach Breeze a voice from your own recordings: between one and thirty minutes of speech in 5 to 900 clips. The trained voice appears in your Voice Library with a Trained badge.

Speed

Measured on an M3 Max, Breeze generates speech in about 0.32 times its playing time (0.38 times with guidance), and playback starts after about 0.29 seconds, while the rest is still being made.

Licence

Breeze TTS 2’s licence is set by its maker. Read it on the Breeze TTS 2 model page before you publish anything you make.

Questions

What is direction?

A short instruction in plain words that shapes how a take is spoken, such as "whispering" or "excited, a bit faster". A Guidance control sets how closely Breeze follows it.

How fast is it?

On an M3 Max, Breeze makes speech in about a third of the time it takes to play (a little longer with guidance on), and the first sound plays after about 0.3 seconds.

Avocado is almost here

Free for Apple silicon Macs running macOS 15 or later. Coming soon.

Coming soon for Mac