[ Notes ]
Notes on speech, timing and subtitles.
Written while building the tool, so everything here is something the code actually does. RSS.
Text to speech with subtitles: the complete guide
How to turn written text into spoken audio and a matching subtitle file in one step — what the formats are, how the timings are produced, where the results are usually wrong, and what to check before you ship.
Read the guide →All notes · 7
Captions for short-form video: TikTok, Reels and Shorts
Short-form captions follow almost none of the broadcast conventions. They are burned in, timed tighter, and read in silence — which changes what a good subtitle file looks like.
Read →Choosing a voice: locale, gender and multilingual
Voice quality is no longer the deciding factor between neural voices. Locale, and whether the voice is multilingual, change the result far more than the name does.
Read →How word-level subtitle timing actually works
Most tools guess subtitle timings from character counts. Speech synthesis can report exactly when each word was spoken — here is what that data looks like and why it changes the result.
Read →SRT or VTT: which subtitle file do you actually need?
The two formats carry the same timings and differ in about four ways. Which one you want depends entirely on where the file is going next.
Read →Subtitle timing rules: cue length, reading speed and line breaks
Perfectly accurate timings still produce unreadable subtitles. The conventions that separate the two are old, well established, and mostly ignored by automatic tools.
Read →Adding subtitles to a YouTube video from a script
If you narrated from a written script, you already have everything YouTube needs. Here is the route that keeps the captions exact instead of letting auto-captions guess.
Read →
