Training builds a small add-on for a voice model from your own recordings, so the model speaks in your voice without needing a reference clip every time. Avocado handles the steps: collect the clips, check them, train, and save the result as a voice.
Collect your clips
Record new samples in Avocado or import recordings you already have. Each clip needs a transcript of what was said; if you have a speech-recognition model installed, Avocado can fill these in for you. A requirements card shows how much speech you have so far and what the model needs.
Train on your Mac
Choose the model to train, the quality you want and press Train. Avocado shows an estimate up front and progress as it goes, and it keeps the best result it found, stopping early if more steps would not help. You can keep using your Mac while it runs.
Use it like any other voice
When training finishes, the voice appears in your Voice Library with a Trained badge. Pick it on the Generate page and use it with the model it was trained on.
Which models can be trained
A dozen model families support training in Avocado, including Breeze TTS 2, Fish Audio S2 Pro, OpenAudio S1 Mini, Orpheus, Qwen3-TTS 0.6B Base, VoxCPM, OuteTTS, MOSS, Dramabox, Irodori TTS and LFM2.5 Audio. Step Audio EditX training is experimental. The Train tab only appears for models that support it.