SMF·COLOSSYAN
Can a prompt replace Colossyan?
AI audio & video generation — synthetic voice, dubbing and AI avatars
Exhibit tracking slip
Verdict
The assembly half is genuinely reachable: a script becomes narration, narration becomes timed slides, slides become an MP4 with captions and a quiz card. That is a local pipeline with ffmpeg and an open TTS model. The presenter is not reachable. Colossyan's avatars are licensed performances of real actors, filmed and cleared for commercial use, driven by a model trained on that footage. You can build the training video; you cannot build the person delivering it.
Exhibit A — The prompt
Received on31.07.2026Build a local training-video generator: a script goes in, a captioned MP4 with slides and narration comes out. No avatar, no face.
Input format: a Markdown script where each level-two heading is a scene. Under a heading, prose becomes narration and a fenced block declares the visual for that scene: a bullet list, an image path, a code sample, or a title card. A scene may also declare a knowledge-check question with options and the correct answer.
Narration: render each scene's prose with a local TTS model (Piper or Kokoro), downloaded on first use. Expose voice, rate and a pronunciation dictionary the user maintains for product names and acronyms, which is the difference between usable and embarrassing in a corporate context. Ship no cloned voices and no celebrity or public-figure voices, and accept only models with a documented licence.
Timing: measure each rendered narration segment and build the timeline from it, rather than asking the user to guess durations. A bullet list reveals one bullet at a time, synchronised to sentence boundaries in the narration. Insert a short pause between scenes.
Rendering: compose slides as SVG from a theme file (fonts, colours, spacing) and rasterise them; assemble video and audio with ffmpeg. Emit a 1080p MP4 plus a WebVTT caption file generated from the narration text with the real timings — captions are a legal requirement for workplace training in many places, and generating them from the source text rather than by re-transcribing is both cheaper and exact.
Knowledge checks: render a declared question as a card at the end of its section with the options visible and a pause, then a card revealing the answer. Also emit the full question set as a JSON file so the same script can feed a quiz in a learning system.
Provenance: write a sidecar JSON per render recording the script hash, TTS model and voice, theme, and output hash. Embed a visible synthetic-media notice in the final frame stating the narration is machine-generated.
Out of scope: any talking-head or avatar generation, cloning a voice, automatic translation into other languages, SCORM or xAPI packaging, and any hosted service.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 the licensed avatars and the rights clearance behind them
- B.2 multi-language versions generated from one script
- B.3 the interactive branching a learning platform expects
- B.4 voice quality close to a professional narrator
Prior art
Exhibit C — Why people still pay: models, compute, rights, and safety operations
Because corporate training video needs a person on screen and a commercial licence for their likeness, and that is a contract with an actor rather than a model you can download.
Questions
Can I import my Colossyan projects?
Scripts can be copied out as text and re-marked-up in Markdown, which is an hour per course. The rendered videos are downloadable and remain usable. Nothing else transfers, since the avatar performances are the platform's own assets.
How does the narration compare to Colossyan's voices?
Noticeably more synthetic. Open TTS models have improved a great deal, but they still flatten emphasis on long sentences and mispronounce domain-specific words unless you build the pronunciation dictionary. For internal training that is often acceptable; for customer-facing material it usually is not.
What does it cost to run?
Nothing after the model downloads. Rendering happens on your own machine at roughly real time or faster, so a ten-minute training video takes a few minutes to produce and costs only electricity.
What is the one thing that does not survive the rebuild?
The presenter. A training video with a person on screen holds attention differently from slides with a voice over them, and Colossyan's avatars come with the model rights and the actor's consent already settled. That is a licensing arrangement, not a model weight.
Related tools
Receipt