Editions
Chatterbox, Chatterbox Multilingual (23 languages), Chatterbox Turbo
Voice cloning
From about 6 to 10 seconds of audio, no transcript needed
Controls
Expressiveness and pacing; nine sound tags on Turbo
Languages
English, or 23 with Multilingual
Memory
About 2 to 2.5 GB
Voice training
Not available
Made by
Resemble AI

Chatterbox is light, quick to clone and easy to steer, which makes it a good everyday choice.

Clone

Each edition has a built-in voice, and any voice you clone can speak with it straight from the recording; no transcript is needed. Turbo wants more than 5 seconds of audio (10 is better); the others do well with about 6. Avocado saves the voice in about a second. Only clone voices you have permission to use.

Expressiveness and pacing

The original and Multilingual editions have two controls: Expressiveness, from flat to dramatic, and Pacing, where lower values are slower and more deliberate.

Turbo tags

Turbo understands nine sounds: laugh, chuckle, sigh, gasp, cough, clear throat, sniff, groan and shush. Pick them from the tag button.

Multilingual

Choose from 23 languages per take. Avocado tells you about each language’s quirks before you generate, for example that Japanese needs kana.

Licence

Chatterbox’s licence is set by Resemble AI. Read it on the Chatterbox model page before you publish anything you make.

Questions

Do I need to type out what was said in the recording?

No. Chatterbox clones from the audio alone.

What's the difference between the editions?

The original is English with expressiveness and pacing controls. Multilingual speaks 23 languages with the same controls. Turbo is faster and adds sound tags like laugh, sigh and gasp.

Avocado is almost here

Free for Apple silicon Macs running macOS 15 or later. Coming soon.

Coming soon for Mac