Lesson 4 of 7 · 4 min

Captions and text on video

Burn word-timed captions onto a video, style them, and understand why captions are the one billed step in the video editor.

Text lanes and burned captions are not the same thing

The video editor gives you two ways to put words on screen, and it helps to keep them apart. Text lanes are for titles, labels, and lower thirds that you type and place by hand. Burned captions are different: they follow the spoken audio, word by word, and you do not type them at all.

How captions get made

To caption a clip, the editor listens to the audio and writes down every word with the exact moment it was said, the way a court stenographer transcribes a hearing, then burns that text onto the video during the stitch. That transcription step runs a model (ElevenLabs speech-to-text), and it is the one billed sub-pass in the entire video editor. Everything else here is free deterministic compute.

Styling captions

Captions do not land on a single plain default. A set of preset styles ships with the editor, and you choose the look before you burn them in.

Re-captioning without paying twice

Because a caption pass costs a model call, the editor keys each one, so asking for the same captions again is not charged a second time. Change the audio and re-caption, though, and that is genuinely new work, so it runs and bills again. If all you need is to fix on-screen wording, use a text lane and skip the caption pass entirely.

Reference

Keep going