wordtospeechEngine ready

Choosing a voice: locale, gender and multilingual

Voice quality is no longer the deciding factor between neural voices. Locale, and whether the voice is multilingual, change the result far more than the name does.

With a few hundred neural voices to pick from, the temptation is to audition names until one sounds nice. That works, but it skips the two choices that actually change the output.

Locale is not cosmetic

A voice identifier looks like en-US-AvaNeural. The middle part is the locale, and it is doing more work than the language alone.

en-GB and en-US are both English. Give an en-GB voice American copy and it will read it correctly — the stress patterns and the rhythm will be right, because those follow the sentence. The vowels will not be. “Schedule”, “privacy”, “route”, every -ile word: they come out British, in copy that reads American, and the mismatch is more noticeable than most people expect before they hear it.

The same applies within a language family far more strongly. es-ES and es-MX are not interchangeable to a listener from either place. Neither are pt-BR and pt-PT, or fr-FR and fr-CA.

Match the locale to your audience, not to the language.

Multilingual voices

Some identifiers contain Multilingualen-US-AvaMultilingualNeural, for instance. These can switch language mid-sentence without changing voice.

Reach for one when your text genuinely mixes languages — a quoted phrase, a product name, a person’s name that should not be anglicised, a bilingual script. Don’t default to one otherwise: for single-language copy, a plain locale voice is usually cleaner, since multilingual models trade away some of their fit to any one locale for the ability to handle all of them.

What gender labels do and do not tell you

The catalogue labels voices male or female, which is useful for narrowing a list of several hundred and not much else. It says nothing about pitch range, pace, warmth, or how the voice handles long-form reading versus short prompts.

Two voices with the same label from the same locale can sound completely different. Audition, do not filter.

Rate and pitch

Both are adjustable, and both are worth adjusting less than you would think.

Rate is genuinely useful. Neural voices often read a touch slower than a person would for the same material, and a small increase — 5 to 10 percent — frequently sounds more natural rather than faster. Beyond about 20 percent the prosody starts to suffer.

Pitch is the one to leave alone unless you have a reason. Shifting it moves the voice away from what the model was trained to produce, and it degrades faster than rate does. If a voice is the wrong pitch for your material, pick a different voice.

One thing worth knowing: changing rate changes the timings. Subtitles generated at 1.0x do not match audio rendered at 1.2x. Set the rate first, then generate — never adjust afterwards and reuse the old subtitle file.

Timing does not vary by voice

A useful fact if you are comparing voices: subtitle accuracy is not one of the variables. Word boundaries come back from the synthesis itself, so every voice produces timings that match its own audio exactly. A slower voice produces longer cues, not worse ones.

Pick on sound. The captions will be right either way.

← All notes