Training builds a small add-on for a voice model from your own recordings, so the model speaks in your voice without needing a reference clip every time. Avocado handles the steps: collect the clips, check them, train, and save the result as a voice.

Collect your clips

Record new samples in Avocado or import recordings you already have. Each clip needs a transcript of what was said; if you have a speech-recognition model installed, Avocado can fill these in for you. A requirements card shows how much speech you have so far and what the model needs.

Train on your Mac

Choose the model to train, the quality you want and press Train. Avocado shows an estimate up front and progress as it goes, and it keeps the best result it found, stopping early if more steps would not help. You can keep using your Mac while it runs.

Use it like any other voice

When training finishes, the voice appears in your Voice Library with a Trained badge. Pick it on the Generate page and use it with the model it was trained on.

Which models can be trained

A dozen model families support training in Avocado, including Breeze TTS 2, Fish Audio S2 Pro, OpenAudio S1 Mini, Orpheus, Qwen3-TTS 0.6B Base, VoxCPM, OuteTTS, MOSS, Dramabox, Irodori TTS and LFM2.5 Audio. Step Audio EditX training is experimental. The Train tab only appears for models that support it.

Questions

How much audio do I need?

For most models, between one minute and half an hour of speech, as anything from 5 to 900 clips of 2 to 30 seconds each. Avocado tracks how much you have and shows what the chosen model needs.

Which models can be trained?

Breeze TTS 2, OpenAudio S1 Mini, Fish Audio S2 Pro, Orpheus, LFM2.5 Audio, OuteTTS, MOSS TTS Local and Nano, Dramabox, Irodori TTS, VoxCPM and Qwen3-TTS 0.6B Base, with Step Audio EditX as an experiment. Avocado only offers the Train tab for models that support it.

How long does it take?

It depends on the model and how many steps you choose. Avocado shows an estimate before you start and progress while it runs, and keeps the best result it found along the way.

Can I train on someone else's recordings?

Only with their permission. Training uses the same rule as cloning: use your own voice or a voice whose owner has agreed.

Avocado is almost here

Free for Apple silicon Macs running macOS 15 or later. Coming soon.

Coming soon for Mac