Lip Sync
Matching a video's mouth movement to a spoken audio track.
Lip sync matches a video subject's mouth movements to a spoken audio track — generating new mouth motion frame by frame rather than relying on footage that was shot with that exact audio. It's the technology behind AI-dubbed video and talking-head avatars generated from a script.
Getting it right depends on tight temporal consistency: every phoneme needs to land on the right frame, or the result reads as uncanny rather than convincing.
See also: Text-to-Speech (TTS), Temporal Consistency.
Put these ideas into motion.
DownloadFrequently asked questions
AI lip sync generates or adjusts a video subject's mouth movements so they match a given audio track, even if the original footage was shot silently or with different words.
Dubbing video into another language, generating talking-head avatars from text, and fixing footage where the audio and mouth movement drifted out of sync during editing.