SMF·WELLSAID-LABS
Can a prompt replace WellSaid Labs?
AI audio & video generation — synthetic voice, dubbing and AI avatars
Exhibit tracking slip
Verdict
WellSaid sells two things a local build separates cleanly. The voices are licensed performances recorded with voice actors under contract, and nothing open matches them for corporate narration. The control surface — adjusting emphasis, pronunciation and pacing on individual words until a take is usable — is reproducible, and it is what turns a passable synthetic voice into one you can ship. Building that editor is a genuine sitting's work and it is where most local TTS setups fall short.
Exhibit A — The prompt
Received on31.07.2026Build a local narration studio whose distinguishing feature is per-word control over delivery.
Model: a local TTS engine — Piper or Kokoro — downloaded on first use. Accept only models with a documented licence and record which licence each voice carries, since that field decides whether the output can be used commercially.
Script: paragraphs of text split into sentences, each sentence rendered as an independent audio segment. Segment-level rendering is what makes iteration bearable: fixing one word re-renders one sentence, not the whole script.
Per-word controls, which are the point of this build. Clicking a word in the script opens controls for that word only:
- Pronunciation override, entered as phonemes in the model's own notation, with a preview button that speaks just that word. Provide a phoneme reference and a lookup that proposes a spelling for an unfamiliar word using espeak-ng's rules as a starting point.
- Emphasis: three levels, applied by adjusting the word's duration and pitch contour relative to its neighbours rather than by shouting it.
- Pause after, in milliseconds.
- Speed multiplier for the word, for slowing a number or an acronym.
Store all of this as a per-word annotation layer above the plain text, so the script stays readable and copyable and the annotations can be exported separately.
Project glossary: a pronunciation dictionary at project level, applied to every occurrence. Adding a word from the per-word editor to the glossary is one click, since the same product name will be wrong in every script otherwise.
The review loop, which matters as much as the controls: waveform of the full narration with sentence boundaries marked, click a sentence to play it, and an A/B toggle between the current and previous render of the same sentence at matched loudness. Matched loudness is essential — otherwise every change sounds like an improvement.
Rendering: assemble sentence segments with the declared pauses, normalise each to a consistent loudness before assembly, and normalise the whole track to a broadcast target. Export WAV and MP3 plus a WebVTT caption file generated from the script with the real segment timings.
Provenance: embed metadata recording the model, voice, licence and render date, and note in the README that synthetic narration used publicly should be disclosed.
Refusals: ship no celebrity, public-figure or scraped voices, and do not implement cloning from a sample.
Out of scope: voice cloning, any hosted API, multi-speaker dialogue, video, and translation.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
Exhibit C — Why people still pay: models, compute, rights, and safety operations
Because a narration voice is a licensed performance. Fifty dollars buys the right to put a specific actor's voice in commercial work, and no open model comes with that.
Questions
Can I bring my WellSaid scripts across?
Scripts are text and copy over trivially. Pronunciation libraries and any per-word tuning are held in WellSaid's own format and do not export, which is the part worth budgeting time to rebuild for recurring terminology.
How large is the voice quality gap really?
Large for long-form narration and smaller for short lines. Open models are intelligible and flat; WellSaid's voices carry prosody across a paragraph in a way that makes a five-minute video comfortable to listen to. Per-word tuning narrows the gap on specific problem words and does not close it.
What does it cost to run?
Nothing per minute after the model downloads. Rendering is faster than real time on a modern machine, so unlimited takes cost only electricity — which is the strongest argument for building this, since iteration is where per-character pricing hurts.
What is the one thing that does not survive the rebuild?
The licence. WellSaid's voices come with the right to use them commercially, backed by contracts with the actors. An open model gives you a sound and, depending on its licence, possibly no right to put it in a client's advertisement at all.
Related tools
Receipt