Back to glossary

Transcription

Converting spoken audio into timestamped text, the basis for captions and subtitles.

Transcription converts spoken audio into text — typically with a timestamp for each word or phrase — turning a voiceover into something searchable and editable as text rather than only as sound.

That per-word timing is also what makes automatic captions possible: once speech is transcribed, the same timestamps place caption text on the timeline in sync with the audio, without typing and timing it by hand.

See also: Text-to-Speech (TTS), Waveform.

Put these ideas into motion.

Download

Frequently asked questions

Transcription converts spoken audio into text, typically with timestamps for each word or phrase, so it can drive captions, subtitles, or a searchable script.

Once speech is transcribed with per-word timing, that timing can be used to auto-place caption text on the timeline in sync with the audio, rather than typing and timing captions by hand.

Ready to tell your story?

Describe an idea and watch the agent animate it — on your Mac, in minutes.