SMF·CAPTIONS-AI
Can a prompt replace Captions?
AI audio & video generation — synthetic voice, dubbing and AI avatars
Exhibit tracking slip
Verdict
Transcribing a video locally with Whisper, burning styled captions onto it with ffmpeg, and cropping to a vertical 9:16 frame using basic face detection to keep the speaker centered is a genuine consolation build — a real one-sitting project that produces a usable captioned clip. It approximates Captions' reframe with a much simpler face-tracking heuristic, not the trained model that keeps a shot framed through fast motion and multiple speakers.
Exhibit A — The prompt
Received on31.07.2026Build a local video-captioning tool: transcribe, burn in styled captions, and crop to vertical with basic face tracking. Stack: Python, ffmpeg for all video processing, whisper.cpp for local transcription, OpenCV for face detection. No cloud API required.
Core loop: the user drops in a local video file. Transcribe it with whisper.cpp to get word-level timestamps, then generate an ASS subtitle track with styled, word-by-word highlighted captions and burn it into the video with ffmpeg. Separately, run OpenCV's face detector on sampled frames, smooth the detected face position across time, and compute a vertical 9:16 crop window that follows it — falling back to a fixed center crop for stretches with no detected face. Export the final captioned, cropped video as a single file.
This needs no API key — whisper.cpp and OpenCV both run fully offline. A GPU speeds up transcription but isn't required for typical short-clip lengths.
Do not build: a trained face/subject-tracking model (this uses off-the-shelf OpenCV detection, explicitly weaker under fast motion or multiple speakers — say so in the UI when confidence is low), AI avatar presenters, or a mobile app. If OpenCV loses the face for more than a couple of seconds, fall back to center crop rather than guessing.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 a trained reframe model that tracks speakers reliably through fast cuts and multiple people in frame
- B.2 a phone-native mobile app for recording and captioning on the go
- B.3 a licensed catalog of AI avatar presenters
- B.4 cloud rendering, so exports don't queue behind your own machine
- B.5 caption style presets tuned and tested across thousands of videos
Prior art
Exhibit C — Why people still pay: models, compute, rights, and safety operations
People pay for Captions because a simple face-detection heuristic loses track of the subject the moment there's fast motion, a hand gesture blocking the face, or more than one person talking — Captions' model was trained specifically to keep tracking through exactly those cases. The subscription buys that reliability under real recording conditions, not just captions and a crop.
Questions
Can I import my existing Captions projects?
No — Captions' project files and caption styles are internal to their app. You'd start from your raw video files and reapply your preferred caption look once as a reusable template.
Will it work on my phone?
No — ffmpeg processing and face detection need more compute than a typical phone browser provides. You'd record on your phone, then run this on a computer to process.
What does it cost to run?
Nothing per video — whisper.cpp and OpenCV both run locally with no per-minute charge. The only cost is the compute time itself, which is free if it's your own machine.
What's the one thing that doesn't survive the rebuild?
Reliable tracking through motion. Captions' model keeps the crop centered on a moving, gesturing speaker or switches correctly between multiple people talking; the OpenCV-based heuristic here loses the face regularly under exactly those conditions and just holds a static crop until it reacquires it.
Related tools
Receipt