SMF·NATURALREADER
Can a prompt replace NaturalReader?
AI audio & video generation — synthetic voice, dubbing and AI avatars
Exhibit tracking slip
Verdict
Plain text-to-speech is a solved local problem. What separates NaturalReader from a simple read-aloud tool is that it handles the documents people actually have: scanned PDFs, photographed pages, screenshots — none of which contain any text until something recognises it. Tesseract does that well and locally, so the OCR-then-speak pipeline is genuinely a sitting's work. The voices are where the gap stays visible, and on a long document that gap gets tiring.
Exhibit A — The prompt
Received on31.07.2026Build a local read-aloud tool that handles scanned documents, not just text.
Input: PDF, EPUB, DOCX, plain text, and images (PNG, JPEG, HEIC).
Text extraction, which is the point of this build. For a PDF, first try the embedded text layer. Detect the scanned case explicitly — a page whose extracted text is empty or is a few stray characters — and fall back to OCR with Tesseract on a rendering of that page at 300 DPI. Handle a mixed document, where some pages have a text layer and some do not, page by page rather than choosing one strategy for the whole file. For images, always OCR.
OCR quality steps that make the difference between usable and useless: deskew the page, convert to greyscale and threshold before recognition, and detect the page's language from a user selection rather than guessing. Two-column layouts must be handled by column detection, because reading a two-column page left to right produces nonsense — this is the single most common failure of naive OCR reading and it should be tested explicitly.
Cleanup before speech: join hyphenated words split across line ends, strip repeated headers and footers detected across pages, drop page numbers, and collapse the line breaks that OCR inserts mid-sentence. Raw OCR text read aloud sounds broken even when the recognition was perfect.
Speech: a local TTS model (Piper for speed, Kokoro for quality), with speed control from 0.5 to 3 times, and a pronunciation dictionary the user maintains. Synthesise a sentence ahead so playback does not stutter at boundaries.
Reading interface: the document text displayed with the sentence being spoken highlighted and auto-scrolled, click any sentence to jump there, and playback position remembered per document. Show the OCR confidence per page and let the user correct recognised text before or during playback, with the correction persisting.
Export: the extracted text as Markdown, and the spoken audio as an MP3 with chapter markers per document section, for listening elsewhere.
Out of scope: mobile apps, any cloud voice, sync between devices, and translation.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
Exhibit C — Why people still pay: models, compute, rights, and safety operations
Because listening to a two-hour document is only tolerable with a voice that does not grate, and voice quality is exactly where hosted models are still meaningfully ahead.
Questions
Is there anything to migrate from NaturalReader?
Your documents are your own files, so nothing is trapped. Reading positions and any custom pronunciation entries do not export, and rebuilding the pronunciation dictionary is the only real setup cost.
How good is the OCR really?
On a clean 300 DPI scan of printed text, Tesseract is excellent. On a phone photo at an angle, or a page with columns, tables and footnotes, it degrades sharply — which is why deskewing, thresholding and column detection are in the prompt rather than left as improvements. Handwriting is out of reach entirely.
What does it cost to run?
Nothing after the models download. OCR and speech both run locally, so an unlimited number of documents costs only processing time — a 300-page scanned book is a few minutes of OCR on a modern machine.
What is the one thing that does not survive the rebuild?
Listening for an hour without getting tired of the voice. Local models are intelligible and flat; the paid voices have prosody that makes long-form listening comfortable, and that is precisely the use case this category exists for.
Related tools
Receipt