File
SMF·D-ID
Received on
31.07.2026
Reviewed on
28.09.2026
Exhibits annexed
3
Questions
4

SMF·D-ID

Can a prompt replace D-ID?

AI audio & video generation — synthetic voice, dubbing and AI avatars

Not yet Verdict recorded on 28.09.2026 · Verified on 31.07.2026
Price
$5.99/moSource: d-id.com · Checked on July 31, 2026
Per year
$71.88
Build time
One sitting
Votes
0 votes
YesAlmostNot yet (checked)

Exhibit tracking slip

Exhibit A The prompt
Exhibit B What you lose
Exhibit C Why people still pay: models, compute, rights, and safety operations
Exhibit Q Questions

Verdict

Open lip-sync models exist and will move a mouth convincingly enough for a short clip, so a local version of the core trick is genuinely reachable. Two things keep this firmly a no. The output quality gap is large — head motion, blinking and the transition at the jawline are what separate a usable clip from an uncanny one, and that is where the proprietary models are ahead. And the consent machinery around a tool that animates faces is not optional, which is exactly the part a personal build has no way to enforce.

Exhibit B — What you lose

Exhibit A — The prompt

Received on31.07.2026
Build a local portrait animation tool with consent and labelling as hard requirements, not settings.

Before any render, the project must record: the source of the portrait, a declaration by the operator that they are the person depicted or hold that person's written permission, the source of the audio, and the intended use. This declaration is stored with the project and reproduced in the output metadata. No render is possible without it.

Refusals, implemented rather than merely documented. Refuse to proceed when the operator declares they do not hold permission. Ship no bundled portraits of real people. Include a prominent statement in the README and in the interface that animating a public figure, or anyone who has not consented, is the thing this tool exists not to do, and that the operator carries the legal and ethical responsibility for every render.

Pipeline: a still portrait plus an audio track. Detect and crop the face, run a local lip-sync model (SadTalker or Wav2Lip weights, downloaded on first use and requiring the user to accept the model's own licence), and composite the animated face back into the original frame with edge blending so the jawline transition is not a hard seam. Emit an MP4 at the source resolution.

Honest quality reporting: after each render, show a side-by-side of the source frame and a sampled output frame, and state in the interface what this class of model does not produce — head movement, natural blinking, and shoulder motion. Do not present the output as indistinguishable from a filmed clip.

Labelling, applied to every output without an opt-out: a visible on-screen notice for the first seconds stating the video is synthetically generated, and C2PA-style provenance metadata embedded in the file recording the source image hash, audio hash, model name and render date. A local provenance log keeps the same record for every render.

Audio: accept a supplied file, or generate narration with a local TTS model under the same rules as the portrait — licensed models only, no cloned or public-figure voices.

Out of scope: real-time or streaming avatars, cloning a voice from a sample, any hosted or API deployment, batch rendering of multiple identities, and removing or weakening the on-screen label. If the user asks for the label to be optional, refuse and explain why.

Opening prefills the prompt — press enter to run it.

Exhibit B — What you lose

  • B.1 the output quality — natural head motion, blinking, a clean jawline
  • B.2 the real-time API behind interactive avatar products
  • B.3 the stock presenter library cleared for commercial use
  • B.4 the consent verification and moderation that make it safe to ship

Prior art

Exhibit C — Why people still pay: models, compute, rights, and safety operations

Because the entire risk sits with whoever renders the video, and a provider that verifies consent, watermarks output and refuses public figures is selling protection as much as pixels.

Questions

Is there anything to migrate from D-ID?

Only the rendered videos, which are downloadable. Source portraits and scripts you supplied are yours already, and D-ID's stock presenters are licensed assets that cannot be exported or reused elsewhere.

How close is the output to D-ID's?

Not close. Local lip-sync models synchronise the mouth reasonably and leave the rest of the face static, which reads as uncanny within a few seconds. D-ID's models generate head movement and blinking, and that is most of the difference between a clip you can publish and one you cannot.

What does it cost to run?

Nothing per render after the model downloads, though a GPU makes the difference between seconds and many minutes per clip. The cost that matters here is not money.

What is the one thing that does not survive the rebuild?

Somebody checking. D-ID verifies consent for likeness use, refuses public figures and watermarks output as policy. A local build enforces whatever the operator declares about themselves, which is a promise rather than a control, and that is precisely why this entry is a no.

Receipt

Already built this yourself?