File
SMF·SPEAK-LANGUAGE
Received on
31.07.2026
Reviewed on
28.09.2026
Exhibits annexed
3
Questions
4

SMF·SPEAK-LANGUAGE

Can a prompt replace Speak?

Learning & languages — AI conversation tutor

Almost Verdict recorded on 28.09.2026 · Verified on 31.07.2026
Price
$17.99/moSource: apps.apple.com · Checked on July 31, 2026
Per year
$215.88
Build time
A weekend
Votes
0 votes
YesAlmost (checked)Not yet

Exhibit tracking slip

Exhibit A The prompt
Exhibit B What you lose
Exhibit C Why people still pay: curriculum and speech feedback
Exhibit Q Questions

Verdict

This one has moved. Speak's core loop — you talk, it understands, it answers in the target language and corrects you — is now assembled from a speech-to-text model, a language model and a voice, all of which are ordinary API calls. That makes it a weekend rather than an impossibility, and the verdict moves off `no` for that reason. What stays behind is the curriculum: Speak decides what you practise and when, and a chat window that will talk about anything is a worse teacher than one with a plan.

Exhibit B — What you lose

Exhibit A — The prompt

Received on31.07.2026
Build a spoken conversation partner for a language you are learning, with the corrections treated as the real output.

The loop: record from the microphone, transcribe with a speech-to-text model, send the transcript plus a short conversation history to a language model, and speak the reply with a text-to-speech voice for the target language. Push-to-talk, not voice activity detection — turn-taking detection is a rabbit hole and a key is a key.

The system prompt does two jobs and they must stay separate. It replies naturally in the target language at a level you configure, and it emits a structured correction block — a JSON array of `{original, corrected, kind, note}` — for anything you said that a native speaker would not. Parse that block out of the response and store it; never read it aloud. A tutor that interrupts every sentence to correct you stops the conversation, which is the one thing you are trying to have.

Give each session a topic and a goal chosen before it starts, from a list you wrote — order food, describe your week, disagree politely. Print the goal at the top and have the model steer towards it. Without that the model drifts into pleasant, useless small talk within four turns.

Afterwards, produce a session review: the transcript, the corrections grouped by kind, and your talking-time ratio. Track corrections across sessions and show the ten most frequent, which is the report that actually makes you better.

Store audio and transcripts locally in dated folders.

Do not claim to score pronunciation. Transcription confidence is not a pronunciation score, and presenting it as one invents a number.

Opening prefills the prompt — press enter to run it.

Exhibit B — What you lose

  • B.1 a designed curriculum, so every session has a point
  • B.2 pronunciation feedback tuned on learner speech rather than generic recognition
  • B.3 the mobile app, where spoken practice actually happens
  • B.4 latency low enough that the conversation feels like one
  • B.5 lessons and progress that work without a connection

Prior art

Exhibit C — Why people still pay: curriculum and speech feedback

Speaking practice fails on friction, and an app that opens to a lesson chosen for you beats a terminal that asks what you would like to talk about.

Questions

Can I bring anything over from Speak?

No — there is no export of your sessions or your error history. Your correction log starts on the first conversation you have with your own build.

How good is the latency?

With separate transcription, generation and speech calls, expect a couple of seconds between your turn and the reply. Realtime speech-to-speech APIs cut that to well under a second at a higher price per minute; both are worth trying, and the slower one is fine for a language you are still assembling sentences in.

What does it cost to run?

A half-hour conversation costs a few tens of cents through separate speech and text calls, more through a realtime speech model. Daily practice lands well under the subscription, which is the point.

What is the one thing that does not survive the rebuild?

Someone deciding what today's lesson is. Speak opens on a chosen scenario at your level; your build opens on a blank prompt, and choosing what to practise is work you now do.

Receipt

Already built this yourself?