Learning strategy

Why speaking practice beats tap-the-word apps

Anyone who has spent six months on a popular streak-based language app and then tried to order a coffee abroad knows the feeling. The words were all there yesterday, on green buttons. Today, in front of a barista in Madrid or Rome, the mouth refuses to open. What happened?

Nothing surprising, once you understand what the app was actually training. The score climbed because it measured what was easy to measure: recognition. The barista is asking for production. These are two different skills, and you can be excellent at the first while remaining a complete beginner at the second.

Recognition is not production

Cognitive science makes an old, boring, well-established distinction between receptive and productive vocabulary. Receptive: you can identify or translate a word when someone shows it to you. Productive: you can retrieve and pronounce that word from scratch when you need it. Native speakers know roughly twice as many words receptively as productively, and the gap for language learners is even wider.

Every tap-the-word interaction is receptive practice. You see four options; you pick the right one. The mental work is recognition — a fundamentally easier task than pulling the word out of empty air, correctly conjugated, mid-sentence, while a stranger waits. Practising recognition for six months mostly makes you better at recognition. It barely moves the production dial.

The three things speaking practice trains that clicking cannot

Why tap-the-word apps got so popular anyway

None of the above is a secret. So why is the app store dominated by exactly the kind of product that does not build the skill people actually want? Three reasons, all economic:

Recognition is gradeable. You clicked the right button or the wrong one. Production is expensive to grade — historically it required a human teacher, and until the last two years, machine grading of spoken language was too shaky to build a product on. So the industry converged on what was easy to score.

Recognition is also rewarding on a five-minute schedule. You get streaks, XP, levelling up. Speaking practice does not offer that dopamine loop as reliably — a real conversation is messy, sometimes embarrassing, and the progress is felt in weeks, not minutes. Products optimising for retention optimise for the loop that keeps users tapping.

Finally, recognition feels like learning. You finish a session, you scored 92%, the streak counter went up. It has the shape of progress even when the underlying skill has barely moved. That is exactly why so many people plateau after a year of daily use and can still not order lunch.

What actually works, and why it is boring

The unglamorous truth is that ten minutes of daily speaking, done consistently for three months, will get an average adult learner to a level that six months of any tap-only app will not. The mechanism is straightforward: you are forcing retrieval, mouth-motor, and recovery in the same session, every session.

Ten minutes is the important number. It is short enough to fit any life, long enough that each session includes several full conversational exchanges. Longer is not obviously better — most learners hit a fatigue wall around twenty minutes of continuous speaking where errors start compounding and the session stops teaching.

The partner does not have to be human. Two years ago the honest answer was italki or a language exchange app; the pool of decent AI roleplay tutors was tiny and the ones that existed did not correct pronunciation or hold a coherent character. That changed. A good AI tutor now beats the median human tutor for early A1–B1 practice on two axes: it is available at 6am for ten minutes, and it has no ego about correcting the same mistake for the eleventh time.

A minimal weekly speaking routine

If you are doing zero speaking practice today and want to fix that this week, here is the smallest routine that works:

  1. Monday, Wednesday, Friday — 10 minutes AI roleplay. Pick a scenario you would plausibly live through (café, train station, small talk with a neighbour). Speak every response out loud, even if you feel ridiculous doing it in your kitchen.
  2. Tuesday, Thursday — 10 minutes shadowing. Play a short native-audio clip (podcast, YouTube monologue, story audio), pause every 15–20 seconds, and repeat what the speaker said as closely as you can. This is the fastest known way to fix a beginner accent.
  3. Saturday — 15 minutes free talk. Longer scenario. No hints, no fallback to English. If the AI (or human) does not understand, you rephrase; you do not switch languages.
  4. Sunday — off, or 10 minutes review. Read your own scrappy notes from the week, out loud. Rest is part of the plan; language consolidation is memory work and memory work needs sleep.

If you insist on keeping the tap-the-word app

There is a case for keeping it as a warm-up. Five minutes of vocabulary drills before a speaking session primes retrieval — recent research on spaced repetition suggests words reviewed just before productive use stick significantly better than words reviewed in isolation. So use the app as an appetizer, not as the meal. The meal is the ten minutes of speaking that comes after.

The other reasonable use of a tap-the-word app is on days when speaking is not possible — you are on public transport, you are sick, you have four minutes between meetings. On those days, recognition practice is genuinely better than nothing. It maintains familiarity even if it does not build production. Just do not confuse maintenance days with growth days, and do not let the streak counter fool you into thinking you have done your work when you have done half of it.

The honest measuring stick

Every three months, do the same test: fifteen minutes of unscripted conversation with a real speaker or a competent AI tutor, no hints, no falling back to English. Record it. Listen to it a week later. Compare it to the recording from three months earlier. That is the only progress metric that matters. Every other number — words known, streak days, app level — is a proxy that can be gamed. The recording cannot.

For the shape of a full plan built around this principle, see the 12-week roadmap to A2 Spanish or the longer roadmap to B2 with an AI tutor. For what each CEFR level actually implies for speaking, the CEFR levels guide is the fastest overview.