Back to glossary

Lip Sync

Matching a video's mouth movement to a spoken audio track.

Lip sync matches a video subject's mouth movements to a spoken audio track — generating new mouth motion frame by frame rather than relying on footage that was shot with that exact audio. It's the technology behind AI-dubbed video and talking-head avatars generated from a script.

Getting it right depends on tight temporal consistency: every phoneme needs to land on the right frame, or the result reads as uncanny rather than convincing.

See also: Text-to-Speech (TTS), Temporal Consistency.

Put these ideas into motion.

Download

Frequently asked questions

AI lip sync generates or adjusts a video subject's mouth movements so they match a given audio track, even if the original footage was shot silently or with different words.

Dubbing video into another language, generating talking-head avatars from text, and fixing footage where the audio and mouth movement drifted out of sync during editing.

Ready to tell your story?

Describe an idea and watch the agent animate it — on your Mac, in minutes.