File
SMF·HEADLINER
Received on
31.07.2026
Reviewed on
28.09.2026
Exhibits annexed
3
Questions
4

SMF·HEADLINER

Can a prompt replace Headliner?

Audio & video editing — podcast recording, cleanup and publishing

Almost Verdict recorded on 28.09.2026 · Verified on 31.07.2026
Price
$19.99/moSource: headliner.app · Checked on July 31, 2026
Per year
$239.88
Build time
A week
Votes
0 votes
YesAlmost (checked)Not yet

Exhibit tracking slip

Exhibit A The prompt
Exhibit B What you lose
Exhibit C Why people still pay: audio infrastructure, distribution, and production polish
Exhibit Q Questions

Verdict

An audiogram is a still image, an animated waveform, and burned-in captions, rendered to video. Every part of that is ffmpeg and a transcript, and a working version is a weekend. Where it stretches is the finishing: caption timing that breaks on phrase boundaries rather than mid-word, a waveform that reads at 400 pixels wide, and a render queue that does not lock up your machine for ten minutes per clip.

Exhibit B — What you lose

Exhibit A — The prompt

Received on31.07.2026
Build a local audiogram generator: audio in, captioned social video out.

Input: an audio or video file. Show a waveform of the whole thing and let the user select a segment by dragging, with fine adjustment by keyboard and a preview of just that segment. Selecting the clip is most of the work in practice, so make it fast.

Transcription: run whisper.cpp locally over the selected segment to get word-level timestamps. Word-level, not segment-level — caption timing depends on it.

Captions, where the quality actually lives. Group words into caption lines using these rules, all configurable:

- A maximum of N characters per line and M lines on screen at once.
- Break at punctuation first, then at a conjunction or preposition, and only then at the character limit. Never break mid-word.
- A minimum on-screen duration per caption, extending short ones by borrowing from the gap that follows.
- Optional word-level highlighting, where the currently spoken word is emphasised within the visible line.

Render every caption from the corrected transcript, and make the transcript editable before rendering — the model will get names and jargon wrong and fixing them afterwards means re-rendering.

Visuals: a background image or solid colour, a title, and an animated waveform driven by the actual audio amplitude, in one of three styles (bars, line, circle). Position and size are chosen from a small set of layouts, not by free placement.

Formats: render the same clip to square (1:1), vertical (9:16) and landscape (16:9) in one pass, with the layout adapting per format and a safe-area guide shown in the preview so captions do not sit under a platform's own interface.

Rendering: compose frames and mux with ffmpeg. Show a real progress bar with an estimated time, run the render in a background process so the interface stays usable, and allow cancellation without leaving a half-written file.

Project files: store the clip selection, corrected transcript, layout and styling as JSON so a clip can be re-rendered after a caption fix without redoing the work.

Out of scope: audio cleanup or mastering, a stock media library, automatic clip selection, direct publishing to any platform, and any hosted rendering.

Opening prefills the prompt — press enter to run it.

Exhibit B — What you lose

  • B.1 cloud rendering — long clips occupy your own machine
  • B.2 the template and stock-media library
  • B.3 automatic suggestions of which segments are worth clipping
  • B.4 the platform presets kept current as social specs change

Prior art

Exhibit C — Why people still pay: audio infrastructure, distribution, and production polish

Because rendering is slow and the presets go stale. A tool that already knows this month's aspect ratio and safe-area rules for four platforms is worth more than the rendering code underneath it.

Questions

Can I import my Headliner projects?

No. Headliner projects live in its own editor with no export beyond the finished video files, which remain usable. Clip selections and caption edits are rebuilt here, which for a back catalogue is not worth doing — start with new episodes.

How long does a render take?

Roughly real time to a few times real time on a modern laptop, depending on resolution and waveform style, so a 60-second clip in three formats is a couple of minutes. That is the fundamental difference from a cloud renderer, and it is why the prompt insists the render runs in the background.

What does it cost to run?

Nothing after the Whisper model downloads. Transcription and rendering are both local, so an unlimited number of clips costs only your machine's time — which is the strongest argument for building this one yourself.

What is the one thing that does not survive the rebuild?

Knowing what to clip. Headliner suggests segments worth pulling out, and choosing the right 45 seconds from an hour is the part that actually determines whether the clip performs. A manual selector leaves that entirely to you.

Receipt

Already built this yourself?