What does parity mean? A detailed comparison of ASR and human transcription errors
Courtney Mansfield, Sara Ng, Gina‐Anne Levow, Mari Ostendorf, RICHARD A. WRIGHT · The Journal of the Acoustical Society of America · 2021
Automatic speech recognition (ASR) has seen dramatic improvements as a result of advances in deep learning. This has led several recent studies to conclude that ASR is approaching parity with human performance, at least in certain speech contexts. These studies use average word error rate (WER) as their primary evaluation metric, dividing the total insertions + substitutions + deletions by the reference length (WER = (I + D+S)/Nr), to compare ASR systems to human transcribers. Averaging combined error types obscures important differences between human and ASR error patterns which impact human comprehension of the output. Human transcribers tend to delete pragmatic and discourse markers (such as fillers and backchannels) and disfluencies, whereas ASR tends to make more substitution errors on words. In conversational settings listeners are able to recover from missing discourse markers, function words, or backchannels, but word substitutions are harder to recover from because they interrupt the information flow. WER is a reasonable first pass metric of ASR performance, but when it comes to communicative parity, it averages out important ways in which it differs from human transcription.