Back to glossary

Voice Cloning

Generating speech in a specific person's voice from a short sample of their real voice.

Voice cloning conditions a text-to-speech model on a sample of a specific person's voice, so new lines can be generated that sound like that person spoke them — rather than a generic, stock voice.

It raises real consent and identity questions that plain TTS doesn't, which is why most voice-cloning tools require verifying rights to the voice being cloned before it can be used.

See also: Text-to-Speech (TTS), Lip Sync.

Put these ideas into motion.

Download

Frequently asked questions

Voice cloning trains or conditions a text-to-speech model on a sample of a specific person's voice, so new speech can be generated that sounds like that person rather than a generic voice.

Voice cloning is a specialized form of text-to-speech — standard TTS uses a stock voice, while cloning targets one specific, identifiable person's voice from a reference sample.

Ready to tell your story?

Describe an idea and watch the agent animate it — on your Mac, in minutes.