SMF·PODSQUEEZE
Can a prompt replace Podsqueeze?
Audio & video editing — podcast recording, cleanup and publishing
Exhibit tracking slip
Verdict
Summarising a transcript is easy and this catalogue already covers it. What makes chapters different is that they need boundaries, not prose: the tool has to find where the conversation actually changed subject and produce a timestamp accurate enough that a listener tapping it lands in the right place. That is a segmentation problem — embedding windows of the transcript and detecting where the topic drifts — and doing it well takes longer than a summary does.
Exhibit A — The prompt
Received on31.07.2026Build a local episode-to-chapters tool. The deliverable is timestamps, not prose.
Transcription: run whisper.cpp locally over the episode with word-level timestamps. Diarise speakers with a simple energy-and-pause heuristic and let the user rename them, propagating through the transcript. Do not claim reliable automatic speaker identification.
Segmentation, which is the core of this build. Chapters come from detecting topic change, not from asking a model to invent them:
1. Split the transcript into overlapping windows of roughly a minute.
2. Embed each window with a local sentence-embedding model.
3. Compute the similarity between consecutive windows and find the local minima — the points where the conversation is least like what preceded it.
4. Filter candidates by a minimum chapter length, configurable and defaulting to two minutes, so a brief digression does not become a chapter.
5. Snap each accepted boundary to the nearest sentence start in the transcript, so a chapter never begins mid-sentence.
6. Only then, ask a model to title each resulting segment from its own text.
This order matters. Asking a model to produce chapters directly from a long transcript yields timestamps that are confidently wrong, because it has no reliable sense of position in a two-hour text. Deriving boundaries mechanically and titling them afterwards produces times a listener can actually trust.
Review: the transcript with detected boundaries marked, each draggable to adjust, with the audio playable from any boundary so the user can hear whether it lands well. Add or remove a boundary by hand and the titles regenerate for the affected segments only.
Outputs, all copy-pasteable:
- Chapters in the format podcast hosts accept: HH:MM:SS followed by the title, one per line.
- A chapters JSON file in the Podcasting 2.0 chapters format.
- Show notes: a paragraph per chapter with its timestamp linked.
- Pull-quotes: passages the model marks as self-contained and quotable, each with its timestamp and speaker so it can be verified against the audio before publishing.
- A full transcript as Markdown and WebVTT.
Use Anthropic or OpenAI for titling and notes, whichever key is configured. With no key, transcription, segmentation and timestamp export must still work — the boundaries are the valuable part and they are computed locally.
Out of scope: audio editing, video clip generation, publishing to any host or CMS, and hosted processing.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 publishing straight into your podcast host or CMS
- B.2 the video clip generation for social
- B.3 the template library for different show formats
- B.4 hosted processing, so a long episode occupies your own machine
Prior art
Exhibit C — Why people still pay: audio infrastructure, distribution, and production polish
Because the output has to land somewhere, and a tool that writes chapters directly into your host's episode form saves the part of the job that is actually tedious.
Questions
Can I import my Podsqueeze outputs?
Show notes and chapters export as text, which is fine to keep as records. There is nothing structured to import, and re-running an episode through this build costs only local processing time.
Why not just ask a model for the chapters directly?
Because it produces plausible titles attached to wrong times. Language models have no reliable sense of position in a long transcript, so the timestamps drift by minutes. Detecting boundaries by embedding similarity gives you positions you can trust, and the model only names what it is given.
What does it cost to run?
Transcription and segmentation are free and local. Only the titling and notes cost anything, and a two-hour episode summarised chapter by chapter is a few cents. A weekly show is under a dollar a year.
What is the one thing that does not survive the rebuild?
The last step. Podsqueeze writes the chapters into your podcast host directly; here you copy a block of text into a form every week. Small, dull, and exactly the kind of friction that makes people keep paying.
Related tools
Receipt