Type the sentence the speaker should read, press Record, read it aloud, press Stop.
For every expected sound (phoneme) in the text, the acoustic model produces a GOP (Goodness of Pronunciation) — how well your audio matches that sound vs. the best competing sound, on the exact milliseconds where you said it.
1. Curve. GOP becomes a 0–100 score through a blend of two curves:
score = (1 − strictness) × calibrated + strictness × strict
The calibrated curve was fitted to human teachers' ratings (forgiving, partial
credit); the strict curve follows the raw acoustic evidence. The slider moves
between them — 0 = teacher-like, 1 = maximally demanding.
2. Variant credit. The profile lists accent variants that are correct for the target accent (e.g. dental t/d for th in Neutral Indian). If your sound matches an accepted variant better than the textbook phone, you get the variant's score (marked ✓ below). The RP profile accepts none.
3. Substitution cap. In strict profiles, a clearly different sound (e.g. w where v belongs) can't score above a ceiling, however confident.
4. Regionalism flags. Known mother-tongue patterns (w-for-v, s-for-sh, j-for-z…) get a ⚑ flag and a tip, so the error is teachable, not just a number.
Word score = average of its phoneme scores (a dropped word-final r after a vowel is optional — British norm — and excluded). Sentence score = average over all phonemes, re-anchored to the human 0–100 scale. Fluency is scored separately from timing (speed, pauses) and is not affected by the accent knobs.