Transcription
Converting spoken audio into timestamped text, the basis for captions and subtitles.
Transcription converts spoken audio into text — typically with a timestamp for each word or phrase — turning a voiceover into something searchable and editable as text rather than only as sound.
That per-word timing is also what makes automatic captions possible: once speech is transcribed, the same timestamps place caption text on the timeline in sync with the audio, without typing and timing it by hand.
See also: Text-to-Speech (TTS), Waveform.
Put these ideas into motion.
DownloadFrequently asked questions
Transcription converts spoken audio into text, typically with timestamps for each word or phrase, so it can drive captions, subtitles, or a searchable script.
Once speech is transcribed with per-word timing, that timing can be used to auto-place caption text on the timeline in sync with the audio, rather than typing and timing captions by hand.