Back to glossary

Text-to-Speech (TTS)

Synthesizing spoken audio from written text using an AI voice model.

Text-to-speech (TTS) converts written text into spoken audio using a trained voice model, producing narration without ever recording a human voice. Modern TTS models render natural pacing, emphasis, and tone rather than the flat, robotic delivery older systems were known for.

GenMotion generates voiceover from a script directly in chat, so narration can go from written line to spoken audio without leaving the editor.

See also: Voice Cloning, Lip Sync.

Put these ideas into motion.

Download

Frequently asked questions

Text-to-speech (TTS) converts written text into spoken audio using a trained voice model, producing narration without recording a human voice.

Yes — GenMotion can generate voiceover from a script directly in chat on the Pro plan, so a scene's narration can be written and read aloud without a separate recording step.

Ready to tell your story?

Describe an idea and watch the agent animate it — on your Mac, in minutes.