SMF·ELEVENLABS
Can a prompt replace ElevenLabs?
AI audio & video generation — voice generation
Exhibit tracking slip
Verdict
A personal text-to-speech tool running a local open-source model is a real weekend build, and it runs entirely offline with no API key. What it cannot responsibly attempt is convincing voice cloning: matching a specific person's voice needs both a much larger model and the consent and safety controls that make cloning something other than a liability.
Exhibit A — The prompt
Received on30.07.2026Build a personal text-to-speech tool around an open-source model that runs locally on CPU (Piper or Kokoro are both good starting points) — paste text or point it at a folder of text files, pick a voice preset, get back audio files, with no API key and no network call required for the core loop. Save every generation with a sidecar text file recording the exact input and the voice used, so nothing is orphaned later. Add an optional cloud fallback: if an API key for a hosted TTS provider is present in the environment, offer it as a higher-quality alternative for a given generation; if the key is missing, skip that option entirely rather than erroring, and the local model keeps working exactly as before.
Do not attempt voice cloning of a specific person, real or fictional — that needs consent, licensing, and safety controls this project cannot provide, and it's exactly the part that should not be a personal weekend build. Be plain in the README about the quality gap against a commercial voice model; local TTS is usable, not indistinguishable from a paid one.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 ElevenLabs' voice quality and range
- B.2 multilingual dubbing
- B.3 safety and consent controls around cloning real voices
- B.4 continuous model updates
Prior art
Exhibit C — Why people still pay: frontier model/safety/licensing
People pay for voices that sound convincingly human and for the licensing and safety work that makes commercial use of a cloned voice defensible — not for text-to-speech itself, which has solid free alternatives.
Questions
Can I bring over voices or audio I generated with ElevenLabs?
The audio files themselves, yes — they're just files. The specific voices are not portable: ElevenLabs' voice models don't export, so a local model gives you a different, less polished voice, not the same one.
Will it work on my phone?
Running the model itself needs real CPU power a phone doesn't have, so this is a desktop or server tool. You could reach a local instance from a phone browser over your own network, but generation happens elsewhere.
What does it actually cost to run?
Nothing, if you stick to the local model — no API key, no hosting fee beyond the machine you already own. The optional cloud fallback costs per character if you turn it on.
What's the one thing that genuinely doesn't survive the rebuild?
Convincing voice cloning. Matching a specific person's voice needs a much larger model plus real consent and safety infrastructure — this build deliberately leaves that out rather than doing it badly.
Related tools
Receipt