wordtospeechEngine ready

Free text to speech with no watermark: what gets added to "free" audio

Audio watermarks, spoken tags, forced title credits and inaudible markers are four different things, and only one of them is what people usually mean by the word.

“No watermark” is one of the most searched qualifiers in this category, and one of the least precise, because at least four different things get called a watermark and they have almost nothing in common.

Sorting them out is worth a minute, because the one people fear most is now rare, and the one that actually affects published work is the one nobody thinks to ask about.

The four things

A spoken tag. A voice reading the tool’s name at the start or end of your audio. Unmissable, impossible to remove cleanly without cutting into your own first or last word, and effectively a demo rather than a product. This is what people picture, and it is much less common than it used to be.

An audible tone or bed. A beep, chime, or low tone laid over the full length of the audio. Rarer still, and generally a sign of a tool built to be upgraded from rather than used.

A forced credit outside the audio. The file itself is clean, but the terms require you to name the tool wherever you publish. Some platforms require it in the title of the work. This is the one that affects real projects most often, and — because the MP3 sounds perfect — it is invisible at exactly the moment you would check for it.

An inaudible marker. A signal embedded in the waveform that you cannot hear but software can detect, used to identify synthetic audio. Increasingly common at the large platforms, driven by provenance and AI-disclosure requirements rather than by branding. It does not affect how your audio sounds or your right to use it, and it is not the same kind of thing as the others at all, despite sharing the word.

That last distinction matters, because a tool can honestly say “no watermark,” meaning nothing audible, while still requiring a credit in your title and still embedding a provenance marker. All three statements are true simultaneously.

What to actually check

Two questions, and the second is the one that gets skipped:

  1. Play the file to the end. A tag at the tail is easy to miss when you only audition the first few seconds — and the end is where they usually sit.
  2. Read what the terms require of your published work. Clean audio and an unconditional right to publish it are separate promises, and a tool can deliver the first while withholding the second.

If a tool advertises “no watermark” but its terms require attribution in your title, you have not really got what you came for. The branding moved from the audio to your headline, where more people see it.

What comes out of this tool

An MP3 containing your text and nothing else. No spoken tag, no tone, no leading or trailing branding, nothing appended at either end.

Nothing has to appear in your title, your description, or your credits. No attribution is required anywhere, and there is no upgrade that removes something — there is nothing being withheld to sell you later.

The subtitle files are the same: SRT and VTT containing your cues, with no inserted credit line as a first or last cue, which is a trick some tools use precisely because people check the audio and not the .srt.

The check worth doing anyway

Whatever tool you use, and regardless of what its page says: export once, listen to the whole file end to end, and open the subtitle file in a text editor before you build anything on top of it.

It takes thirty seconds, it is the only way to actually know rather than assume, and it is dramatically cheaper than discovering a trailing voice tag after you have cut twenty clips around the audio.

← All notes