File
SMF·PLAYHT
Received on
31.07.2026
Reviewed on
28.09.2026
Exhibits annexed
3
Questions
4

SMF·PLAYHT

Can a prompt replace PlayHT?

AI audio & video generation — synthetic voice, dubbing and AI avatars

Not yet Verdict recorded on 28.09.2026 · Verified on 31.07.2026
Price
$39/moSource: play.ht · Checked on July 31, 2026
Per year
$468
Build time
One sitting
Votes
0 votes
YesAlmostNot yet (checked)

Exhibit tracking slip

Exhibit A The prompt
Exhibit B What you lose
Exhibit C Why people still pay: models, compute, rights, and safety operations
Exhibit Q Questions

Verdict

You can stand up a local TTS server with a streaming endpoint in an afternoon, and for reading a script it will do. What PlayHT sells is the number that matters to a voice agent: time to first audio byte, held under a few hundred milliseconds, from a fleet sized so it stays there under load. A local model on one machine cannot hold that budget while a language model is also generating the words, and the voice catalogue is licensed, not downloadable.

Exhibit B — What you lose

Exhibit A — The prompt

Received on31.07.2026
Build a local streaming text-to-speech server with a latency budget you can measure.

Model: run an open-weights TTS model locally — Piper for speed and small footprint, or Kokoro for better quality at higher cost. Make the model a swappable module behind one interface and ship both, downloaded on first use.

Streaming endpoint: POST text, receive audio as a chunked HTTP stream, first chunk emitted as soon as the first sentence is synthesised rather than after the whole input. Split incoming text on sentence boundaries and synthesise sentence by sentence into a queue, so a long paragraph starts playing immediately. Support both raw PCM and Opus output; Opus for anything crossing a network.

Latency instrumentation, which is the point of this build: measure and expose time to first audio byte, real-time factor (audio seconds produced per wall-clock second), and queue depth. Log every request with those three numbers. Provide a benchmark command that fires N concurrent requests and prints the distribution, so the user can see exactly where their machine stops keeping up. Do not hide these numbers behind a debug flag — they are the reason to run this rather than call an API.

Voices: expose the model's built-in voices with speed and pitch controls, and a pronunciation dictionary the user maintains for names and jargon the model gets wrong. Ship no cloned voices and no celebrity or public-figure voices, and refuse to load a voice model that does not come with a documented licence — state this in the README.

Caching: hash the text plus voice plus settings and cache the rendered audio on disk, so repeated phrases (a greeting, an error message) return instantly. For an agent workload this is where the real latency win comes from.

Out of scope: voice cloning, any hosted or multi-tenant deployment, an account system, and a timeline or editing interface. This is a server, not a studio.

Opening prefills the prompt — press enter to run it.

Exhibit B — What you lose

  • B.1 sub-300-millisecond time to first byte under concurrent load
  • B.2 the licensed voice catalogue, including the ones cleared for commercial use
  • B.3 voice cloning with the consent verification that makes it usable commercially
  • B.4 capacity — one machine serves one or two concurrent streams, not a product

Prior art

Exhibit C — Why people still pay: models, compute, rights, and safety operations

Because a voice agent that pauses for a second before answering sounds broken, and closing that gap is a GPU fleet and a latency-tuned serving stack, not a model download.

Questions

Can I use my PlayHT voices with this?

No. PlayHT's voices are licensed assets served from its own infrastructure and there is nothing to export. You get whatever voices the open-weights model ships with, and they sound noticeably more synthetic than PlayHT's better ones.

What latency can I actually expect?

On a modern laptop CPU, Piper reaches first audio in roughly 100 to 200 milliseconds for a short sentence, which is genuinely competitive. Kokoro is better sounding and several times slower. The gap opens under concurrency: PlayHT holds its latency across many simultaneous streams and one machine does not.

What does it cost to run?

Nothing per character, which inverts the economics entirely. PlayHT bills by usage; this bills by whatever machine you already have. For a personal project generating a lot of speech, that difference is the strongest argument for building it.

What is the one thing that does not survive the rebuild?

Concurrency. A single-machine server handles one or two live streams before latency collapses, so this is fine behind a personal assistant and unusable behind anything with real users. Scaling it means renting GPUs, which is the business PlayHT is in.

Receipt

Already built this yourself?