SMF·ELSA-SPEAK
Can a prompt replace ELSA Speak?
Learning & languages — English pronunciation coaching
Exhibit tracking slip
Verdict
ELSA reports which phoneme you got wrong, not just whether you were understood, and that requires a model trained on a large corpus of non-native speech labelled by phoneticians. Speech recognition is available to anyone now; labelled learner speech is not, and without it the best a personal build can do is tell you whether a recogniser heard the right word.
Exhibit A — The prompt
Received on31.07.2026Build a local pronunciation practice loop for a language you are learning, and be precise about what it can and cannot tell you.
The loop: display a target sentence, record from the microphone, transcribe locally with whisper.cpp, and compare the transcript to the target with a word-level alignment. Show each word coloured by whether it was recognised, substituted or dropped, and keep the audio file beside the attempt so you can listen back.
Be honest in the interface about what this measures. A word the recogniser missed is evidence that something is off; a word it caught is not evidence that you said it well, because recognisers are trained to be forgiving of accent. Put that sentence in the UI, not in a README, because the temptation to read the output as a pronunciation score is strong and wrong.
Run each attempt through the recogniser three times with different temperature settings and report agreement across runs. A word that is transcribed differently on each pass is genuinely unclear in your recording — that instability is a better signal than any single transcript.
Build the error report over time: words dropped or substituted most often across all your attempts, with links to the recordings. That list is what you take to a tutor.
Do not compute a phoneme-level score. Forced alignment tools exist and will happily produce numbers, but calibrating them against native speech is the actual research problem here, and an uncalibrated score is a confident number about nothing.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 per-phoneme scoring, which is the entire product
- B.2 targeted drills chosen from your own error profile
- B.3 the curriculum of sentences graded by difficulty
- B.4 the mobile app and its recording workflow
Prior art
- WhisperLicense: MIT
Exhibit C — Why people still pay: pronunciation model trained on learner speech
Being told that your /θ/ is landing as /s/ is actionable in a way that "the computer misheard you" never is, and no open model produces that judgement.
Questions
Can I get my ELSA history out?
No. ELSA exposes no export of your recordings or your per-sound scores, so the practice history stays inside the app. Your own log starts empty and grows from your recordings.
Does whisper.cpp run fast enough on a laptop?
For single sentences, yes. A small or medium model transcribes a few seconds of audio in well under a second on modern hardware, and the whole loop stays offline.
What does it cost to run?
Nothing. Local model, local audio, no key. Model weights are a one-off download of a few hundred megabytes.
What is the one thing that does not survive the rebuild?
Being told which sound was wrong. Your version can say a word did not land; ELSA says the vowel was too far forward, and that is the difference between knowing you have a problem and knowing how to fix it.
Related tools
Receipt