wordtospeechEngine ready

Captions for short-form video: TikTok, Reels and Shorts

Short-form captions follow almost none of the broadcast conventions. They are burned in, timed tighter, and read in silence — which changes what a good subtitle file looks like.

Nearly everything written about subtitle conventions assumes broadcast: a sixteen-by-nine frame, a viewer with the sound on, two lines at the bottom. Short-form video breaks all three assumptions, and captions built to broadcast rules look wrong there.

What is different

Most viewers have the sound off, so captions aren’t an accessibility add-on on these platforms — they’re the primary channel, and a clip without them gets scrolled past before the first sentence lands. They’re burned in rather than attached, too: platform caption tracks exist, but convention renders captions into the video itself, styled and positioned toward the middle of the frame instead of the bottom.

The frame being vertical matters more than it sounds. A 42-character line that sits comfortably in landscape wraps awkwardly at the text sizes these platforms use. And attention is measured in the first second — captions need to appear immediately, not after a beat, because that’s usually all the time a clip gets before someone decides whether to keep watching.

What that changes

Cues get shorter. Where broadcast tolerates 42 characters over two lines, short-form wants roughly half that — often one phrase at a time, sometimes just a word or two for emphasis. They change faster and each one holds less.

Timing gets tighter as well. A caption arriving 300 milliseconds late is invisible in a five-minute video and obvious in a fifteen-second one, which is exactly where estimated timings fall apart fastest — there’s no room left for drift.

How word-level subtitle timing actually works

And captions belong in the centre, not the bottom. Platform UI — usernames, captions, buttons — occupies the lower third, so text placed there out of broadcast habit ends up sitting underneath the interface.

The workflow that works

1. Write the script first. Short-form rewards tight writing anyway, and a script means the caption text is exact rather than transcribed.

2. Generate audio and timings together. If you are narrating with a synthesised voice, the engine can report when each word was spoken, so the caption timing is measured rather than estimated.

Text to SRT: from a paragraph to a subtitle file

3. Import the SRT into your editor. CapCut, Premiere, Final Cut and the rest all accept SRT and will convert cues into styled text layers you can restyle in bulk.

SRT or VTT: which subtitle file do you actually need?

4. Split the cues down. This is the manual step. Cues sized for broadcast need breaking into phrases. Most editors let you split a caption at the playhead, which is fast once the timings underneath are correct.

5. Upload the SRT as well. Even with captions burned in, attaching the file gives you searchable text and a translation source. It costs nothing.

The one rule that carries over

Read a cue out loud. If you cannot finish it comfortably in the time it is on screen, it is too long — and in a vertical frame at short-form pace, “too long” arrives much sooner than the broadcast numbers suggest.

Subtitle timing rules: cue length, reading speed and line breaks

← All notes