🎤 ZeneSpeech — Test Console

Type the sentence the speaker should read, press Record, read it aloud, press Stop.

How the score is calculated (and what the knobs do)

For every expected sound (phoneme) in the text, the acoustic model produces a GOP (Goodness of Pronunciation) — how well your audio matches that sound vs. the best competing sound, on the exact milliseconds where you said it.

1. Curve. GOP becomes a 0–100 score through a blend of two curves:
  score = (1 − strictness) × calibrated + strictness × strict
The calibrated curve was fitted to human teachers' ratings (forgiving, partial credit); the strict curve follows the raw acoustic evidence. The slider moves between them — 0 = teacher-like, 1 = maximally demanding.

2. Variant credit. The profile lists accent variants that are correct for the target accent (e.g. dental t/d for th in Neutral Indian). If your sound matches an accepted variant better than the textbook phone, you get the variant's score (marked below). The RP profile accepts none.

3. Substitution cap. In strict profiles, a clearly different sound (e.g. w where v belongs) can't score above a ceiling, however confident.

4. Regionalism flags. Known mother-tongue patterns (w-for-v, s-for-sh, j-for-z…) get a ⚑ flag and a tip, so the error is teachable, not just a number.

Word score = average of its phoneme scores (a dropped word-final r after a vowel is optional — British norm — and excluded). Sentence score = average over all phonemes, re-anchored to the human 0–100 scale. Fluency is scored separately from timing (speed, pauses) and is not affected by the accent knobs.