Ask a room of North Americans to say cot and then caught, and the room will quietly split in two. Roughly half will say the words identically, both with the open “ah” of father (/ɑː/). The other half will keep caught on a separate, rounder vowel — the “aw” of law (/ɔː/) — and may be genuinely surprised to learn their neighbours do not. Neither half is wrong. This is the cot–caught merger, one of the great quiet divides of American speech, and it matters more to pronunciation apps than almost anyone admits.
A vowel merger is what happens when two historically separate vowels drift together until a community of speakers no longer distinguishes them. For merged speakers, cot and caught, stock and stalk, don and dawn are perfect homophones — the distinction simply is not part of their English, any more than the distinctions of Chaucer’s English are part of yours. For unmerged speakers, the pairs remain crisply different, and the “aw” vowel is a fully paid-up member of their inventory.
Both varieties are correct General American. The merger is geography and generation, not carelessness: it is not a sound one group failed to learn, but a contrast one group’s English no longer carries. And crucially, nothing goes wrong when the two halves talk to each other. Context absorbs the ambiguity so completely that most people live decades without noticing which side of the line they are on.
Now consider what happens when a pronunciation app hard-codes a single dictionary answer for these words. If the reference recording keeps “caught” on the “aw” vowel, then every merged speaker — half a continent of native and near-native voices — gets marked down for producing perfectly good American English. If the reference merges them, the unmerged half is penalised instead. Either way the app is punishing geography and calling it accuracy.
The deeper point is about what a diagnosis is for. A pronunciation coach is only useful if it flags differences that carry meaning — the substitutions that make listeners hear a different word, or strain to hear the intended one. A learner who says “ship” with the vowel of “sheep” has one of those. A speaker who says “caught” the Californian way rather than the New York way has nothing of the kind; they have an address. A scoring system that cannot tell those two situations apart is not strict — it is confused.
“In my village, the next town over sang the same aria with different vowels, and both broke your heart. You correct the note that changes the meaning — the rest is scenery!” — Nunzio, of course, has opinions.
Nunzio’s scoring treats cot–caught pairs as accepted variants: whichever of the two vowels you produce in these words, the app scores it as right. Merged speakers are never told to un-merge; unmerged speakers are never told their “aw” is an error. At the same time, the two vowels remain distinct entries in the sound library — the “ah” of father and the “aw” of law each get their own page — because unmerged speakers need both, and because plenty of words are not covered by the merger at all.
That combination — distinct in the library, interchangeable where the dialects genuinely disagree — is what it means to score pronunciation honestly. The alternative, one canonical answer per word, is easier to build and quietly wrong for millions of speakers.
The Nunzio app scores exactly these sounds on your own phone, phoneme by phoneme — and where American English itself accepts two answers, so does the score.