File
SMF·SWELL-AI
Received on
31.07.2026
Reviewed on
28.09.2026
Exhibits annexed
3
Questions
4

SMF·SWELL-AI

Can a prompt replace Swell AI?

Audio & video editing — podcast recording, cleanup and publishing

Almost Verdict recorded on 28.09.2026 · Verified on 31.07.2026
Price
$29/moSource: www.swellai.com · Checked on July 31, 2026
Per year
$348
Build time
A week
Votes
0 votes
YesAlmost (checked)Not yet

Exhibit tracking slip

Exhibit A The prompt
Exhibit B What you lose
Exhibit C Why people still pay: audio infrastructure, distribution, and production polish
Exhibit Q Questions

Verdict

Audio is invisible to search. Swell's real job is converting an episode into on-page text with the structure that makes it findable: a full transcript with speaker labels and anchors, chapter headings, and the structured data that tells a search engine what the page is. That is a transcription pass plus careful HTML generation, which is a solid few days. The part you keep paying for is that Swell publishes it into your host or CMS, which is the step nobody enjoys doing by hand every week.

Exhibit B — What you lose

Exhibit A — The prompt

Received on31.07.2026
Build a tool that turns a podcast episode into a publishable transcript page. The output is structured HTML, not a blog post.

Transcription: whisper.cpp locally with word-level timestamps. Speaker splitting by an energy-and-pause heuristic, with the user renaming speakers once and the change propagating. Do not claim automatic speaker identification.

Structuring, which is what makes the page useful rather than a wall of text:

- Merge word-level output into speaker turns, then into paragraphs at natural pauses, so the transcript reads as dialogue rather than a stream.
- Detect chapter boundaries by embedding consecutive windows of the transcript and finding topic shifts, then snap each boundary to a turn start. Title each chapter from its own text.
- Give every chapter an anchor id and every paragraph a timestamp attribute, so a link can point at a moment.
- Produce an episode summary and a short list of key points from the transcript.

The page output, which is the deliverable:

- Semantic HTML: an article element, chapter headings as h2 with anchors, speaker turns marked up consistently, timestamps rendered as links that also carry a machine-readable time attribute.
- JSON-LD structured data using the PodcastEpisode type, with the episode name, description, duration, publication date, the associated series, and the audio file URL. This is the single highest-value part of the output and it must be valid — validate it in the test suite rather than assuming.
- A table of contents built from the chapter anchors.
- Meta title and description, with lengths checked against the usual truncation points.
- An audio player element pointing at the episode file, with the transcript below it.

Editing before export: the transcript is editable in the tool, because Whisper will get names, jargon and numbers wrong and publishing those is worse than not publishing. A per-episode glossary of correct spellings, applied as a find-and-replace pass on every future episode of the same show, so the same names stop being wrong every week.

Exports: the standalone HTML page, a Markdown version for a static site generator, plain WebVTT captions for the audio player, and the JSON-LD block on its own for pasting into a CMS.

Honesty about what this does not do: it produces a page, it does not publish one, and it makes no claim about ranking. Say both in the README.

Use Anthropic or OpenAI for chapter titles and the summary. With no key, transcription, structuring, anchors and the HTML export must all still work — only titles and the summary are unavailable.

Out of scope: publishing to any host or CMS, audio editing, video clip generation, and social post generation.

Opening prefills the prompt — press enter to run it.

Exhibit B — What you lose

  • B.1 publishing straight into your podcast host or CMS
  • B.2 hosted processing, so long episodes occupy your own machine
  • B.3 the template library for different show formats
  • B.4 the social clip generation

Prior art

Exhibit C — Why people still pay: audio infrastructure, distribution, and production polish

Because the value only lands once the page is live, and the last step — getting it into the CMS every week without fuss — is the one a local tool leaves you doing.

Questions

Can I import my Swell content?

Transcripts and generated text export and can be kept as records. There is nothing structured to import, and re-processing an episode costs only local time, so a back catalogue can simply be re-run.

Does a transcript page actually help discovery?

It gives search engines text where there was none, which is a precondition rather than a guarantee. The structured data helps a search engine understand what the page is. Neither makes an episode rank on its own, and any tool claiming otherwise is overselling.

What does it cost to run?

Transcription is free and local. Chapter titles and the summary are a few cents per episode. A weekly show costs a couple of dollars a year, against seventeen dollars a month.

What is the one thing that does not survive the rebuild?

The publish step. Swell writes the page into your CMS; here you export HTML and paste it in every week. Small, repetitive, and exactly the friction that eventually stops people doing it at all.

Receipt

Already built this yourself?