Speakers
Up to 4 per script
Languages
English and Chinese
Voice cloning
From audio alone; 10 to 30 s works best
Long-form
Speaks the whole script in one pass
Playback
Plays while it speaks
Memory
About 7.5 GB at peak
Made by
Microsoft
A two-speaker dialogue with VibeVoice 1.5B, with Ryan and Serena cast as speakers 1 and 2

VibeVoice 1.5B from Microsoft is made for long-form speech such as podcasts and conversations, where voices need to stay the same from the first line to the last.

Write the script

Put each turn on its own line, starting with Speaker 1: up to Speaker 4:. The tag button inserts these for you. Plain text is spoken by one speaker.

Cast voices

Press Dialogue and choose a voice for each speaker. Cloning needs only audio, no transcript; 10 to 30 seconds works best. Only clone voices you have permission to use.

Quality

A Quality setting trades speed for detail, and Voice likeness sets how closely a take follows the cast voice.

Licence

VibeVoice’s licence and usage notes are set by Microsoft. Read them on the VibeVoice 1.5B model page before you publish anything you make.

Questions

What about VibeVoice Realtime?

Avocado also works with VibeVoice Realtime, a smaller edition with 25 preset voices and near-instant playback. It doesn't clone.

Do I have to cast every speaker?

No. Without a cast, VibeVoice picks a voice for each speaker. Pin the seed to keep the same ones.

Avocado is almost here

Free for Apple silicon Macs running macOS 15 or later. Coming soon.

Coming soon for Mac