Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment
Ethan A. Chi, Julián Salazar, Katrin Kirchhoff · 2021
Non-autoregressive encoder-decoder models greatly improve decoding speed over autoregressive models, at the expense of generation quality.To mitigate this, iterative decoding models repeatedly infill or refine the proposal of a non-autoregressive model.However, editing at the level of output sequences limits model flexibility.We instead propose iterative realignment, which by refining latent alignments allows more flexible edits in fewer steps.Our model, Align-Refine, is an end-to-end Transformer which iteratively realigns connectionist temporal classification (CTC) alignments.On the WSJ dataset, Align-Refine matches an autoregressive baseline with a 14× decoding speedup; on LibriSpeech, we reach an LM-free testother WER of 9.0% (19% relative improvement on comparable work) in three iterations.We release our code at https://github.com/ amazon-research/align-refine.