File
SMF·LOVO
Received on
31.07.2026
Reviewed on
28.09.2026
Exhibits annexed
3
Questions
4

SMF·LOVO

Can a prompt replace LOVO?

AI audio & video generation — synthetic voice, dubbing and AI avatars

Not yet Verdict recorded on 28.09.2026 · Verified on 31.07.2026
Price
$24/moSource: lovo.ai · Checked on July 31, 2026
Per year
$288
Build time
One sitting
Votes
0 votes
YesAlmostNot yet (checked)

Exhibit tracking slip

Exhibit A The prompt
Exhibit B What you lose
Exhibit C Why people still pay: models, compute, rights, and safety operations
Exhibit Q Questions

Verdict

The studio around the model is buildable: a script where each line has a speaker, an emphasis marker and a pause, rendered line by line and mixed into one timeline. Open TTS models handle the rendering. What Lovo sells is the catalogue — hundreds of voices, in dozens of languages, licensed for commercial use, with emotion variants recorded or trained per voice. Open weights give you a handful of voices with one register each, and no rights paperwork behind them.

Exhibit B — What you lose

Exhibit A — The prompt

Received on31.07.2026
Build a local multi-speaker voiceover studio. A script goes in, one mixed audio file comes out.

Script format: plain text where a line begins with a speaker name and a colon. Blank lines are pauses. Inline markers control delivery: a bracketed number for an explicit pause in milliseconds, and a marker for emphasis on a word or phrase. Parse errors are reported with a line number, never guessed at.

Speakers: each maps to a voice from a locally installed TTS model, with per-speaker rate, pitch offset and volume. Accept only models with a documented licence, and record which licence each speaker's voice carries in the project file — this is the field that decides whether the output can be used commercially, and it should be visible rather than buried.

Rendering: synthesise each line separately, not the whole script at once. Line-by-line rendering means one bad line can be re-rendered without redoing the take, and it is what makes iteration bearable. Cache each line by a hash of its text plus voice plus settings, so re-rendering after editing one line costs one line.

Mixing: assemble the rendered lines on a timeline with the declared pauses between them, plus a small configurable default gap. Normalise each line to a consistent loudness before mixing so a quiet voice does not disappear next to a loud one, then normalise the whole track to a target LUFS. Optional background music bed at a fixed duck level under speech, from a file the user supplies — ship no music.

Pronunciation: a project-level dictionary mapping written forms to phonetic spellings, applied before synthesis. This is the difference between usable and embarrassing for product names, and it should be a first-class part of the interface, not a settings page.

Review: a waveform of the mixed track with line boundaries marked and the script beside it, so clicking a line plays it and clicking re-render replaces just that line in the mix.

Provenance and labelling: embed metadata recording the models and voices used, their licences, and the render date. Write a sidecar file with the same. Note in the README that synthetic speech used publicly should be disclosed.

Refusals: ship no celebrity, public-figure or scraped voices, and do not implement voice cloning from a sample.

Out of scope: voice cloning, any hosted API or account, video, and automatic translation of the script.

Opening prefills the prompt — press enter to run it.

Exhibit B — What you lose

  • B.1 the voice catalogue and its commercial licensing
  • B.2 emotion and delivery variants per voice
  • B.3 the language coverage that makes a dubbing workflow possible
  • B.4 cloning a voice with the consent workflow that makes it usable

Prior art

Exhibit C — Why people still pay: models, compute, rights, and safety operations

Because a voiceover is a commercial asset and its voice needs a licence attached. Open weights get you a sound; they do not get you the right to put it in a client's advertisement.

Questions

Can I bring my Lovo projects over?

Scripts can be copied out as text and re-marked-up, which is quick. The voices cannot: they are Lovo's licensed assets, so every speaker is re-cast from whatever local models you have, and the result will sound different.

Can I use the output commercially?

Only if the voice model's licence permits it, which is why the prompt records the licence per speaker in the project file. Some open TTS weights are permissively licensed and some are research-only, and the distinction is easy to lose track of once a project has five speakers.

What does it cost to run?

Nothing per minute after the models download. Rendering happens locally at faster than real time on a modern machine, so a ten-minute script takes a couple of minutes and costs only electricity — against per-minute or credit-based pricing on the paid product.

What is the one thing that does not survive the rebuild?

Casting. Lovo's value is browsing hundreds of voices until one fits the piece, in the language you need, with the emotional register you want. A handful of local voices means you write the script to suit the voices you have.

Receipt

Already built this yourself?