File
SMF·WELLSAID-LABS
Received on
31.07.2026
Reviewed on
28.09.2026
Exhibits annexed
3
Questions
4

SMF·WELLSAID-LABS

Can a prompt replace WellSaid Labs?

AI audio & video generation — synthetic voice, dubbing and AI avatars

Not yet Verdict recorded on 28.09.2026 · Verified on 31.07.2026
Price
$50/moSource: wellsaidlabs.com · Checked on July 31, 2026
Per year
$600
Build time
One sitting
Votes
0 votes
YesAlmostNot yet (checked)

Exhibit tracking slip

Exhibit A The prompt
Exhibit B What you lose
Exhibit C Why people still pay: models, compute, rights, and safety operations
Exhibit Q Questions

Verdict

WellSaid sells two things a local build separates cleanly. The voices are licensed performances recorded with voice actors under contract, and nothing open matches them for corporate narration. The control surface — adjusting emphasis, pronunciation and pacing on individual words until a take is usable — is reproducible, and it is what turns a passable synthetic voice into one you can ship. Building that editor is a genuine sitting's work and it is where most local TTS setups fall short.

Exhibit B — What you lose

Exhibit A — The prompt

Received on31.07.2026
Build a local narration studio whose distinguishing feature is per-word control over delivery.

Model: a local TTS engine — Piper or Kokoro — downloaded on first use. Accept only models with a documented licence and record which licence each voice carries, since that field decides whether the output can be used commercially.

Script: paragraphs of text split into sentences, each sentence rendered as an independent audio segment. Segment-level rendering is what makes iteration bearable: fixing one word re-renders one sentence, not the whole script.

Per-word controls, which are the point of this build. Clicking a word in the script opens controls for that word only:

- Pronunciation override, entered as phonemes in the model's own notation, with a preview button that speaks just that word. Provide a phoneme reference and a lookup that proposes a spelling for an unfamiliar word using espeak-ng's rules as a starting point.
- Emphasis: three levels, applied by adjusting the word's duration and pitch contour relative to its neighbours rather than by shouting it.
- Pause after, in milliseconds.
- Speed multiplier for the word, for slowing a number or an acronym.

Store all of this as a per-word annotation layer above the plain text, so the script stays readable and copyable and the annotations can be exported separately.

Project glossary: a pronunciation dictionary at project level, applied to every occurrence. Adding a word from the per-word editor to the glossary is one click, since the same product name will be wrong in every script otherwise.

The review loop, which matters as much as the controls: waveform of the full narration with sentence boundaries marked, click a sentence to play it, and an A/B toggle between the current and previous render of the same sentence at matched loudness. Matched loudness is essential — otherwise every change sounds like an improvement.

Rendering: assemble sentence segments with the declared pauses, normalise each to a consistent loudness before assembly, and normalise the whole track to a broadcast target. Export WAV and MP3 plus a WebVTT caption file generated from the script with the real segment timings.

Provenance: embed metadata recording the model, voice, licence and render date, and note in the README that synthetic narration used publicly should be disclosed.

Refusals: ship no celebrity, public-figure or scraped voices, and do not implement cloning from a sample.

Out of scope: voice cloning, any hosted API, multi-speaker dialogue, video, and translation.

Opening prefills the prompt — press enter to run it.

Exhibit B — What you lose

  • B.1 the voice catalogue and the actor contracts behind it
  • B.2 the commercial rights that let the output go in a client's advertisement
  • B.3 the voice quality that makes long narration comfortable to listen to
  • B.4 team workflows and shared pronunciation libraries

Prior art

Exhibit C — Why people still pay: models, compute, rights, and safety operations

Because a narration voice is a licensed performance. Fifty dollars buys the right to put a specific actor's voice in commercial work, and no open model comes with that.

Questions

Can I bring my WellSaid scripts across?

Scripts are text and copy over trivially. Pronunciation libraries and any per-word tuning are held in WellSaid's own format and do not export, which is the part worth budgeting time to rebuild for recurring terminology.

How large is the voice quality gap really?

Large for long-form narration and smaller for short lines. Open models are intelligible and flat; WellSaid's voices carry prosody across a paragraph in a way that makes a five-minute video comfortable to listen to. Per-word tuning narrows the gap on specific problem words and does not close it.

What does it cost to run?

Nothing per minute after the model downloads. Rendering is faster than real time on a modern machine, so unlimited takes cost only electricity — which is the strongest argument for building this, since iteration is where per-character pricing hurts.

What is the one thing that does not survive the rebuild?

The licence. WellSaid's voices come with the right to use them commercially, backed by contracts with the actors. An open model gives you a sound and, depending on its licence, possibly no right to put it in a client's advertisement at all.

Receipt

Already built this yourself?