Evaluation of Speech Translation Subtitles Generated by ASR with Unnecessary Word Detection
Makoto Hotta, Chee Siang Leow, Norihide Kitaoka, Hiromitsu Nishizaki · 2024
This study addresses the problem of generating understandable speech translation subtitles for spontaneous speech, such as lectures and talks, which often contain disfluencies like fillers, voiced pauses, and rephrased utterances. We propose a speech translation method that considers (1) the translation unit (speech unit or sentence unit) input to the machine translation system and (2) the use of an end-to-end automatic speech recognition (ASR) system that detects unnecessary words for translation while recognizing speech. A subtitle evaluation experiment with subjects who do not understand Japanese showed that machine translation using ASR transcriptions automatically removed unnecessary words and considered the translation unit could produce better subtitles. The results suggest that translating sentence by sentence makes the generated subtitle sentences more comprehensible, and using an ASR model with unnecessary word detection is particularly effective for generating easy-to-understand subtitles in sentence-by-sentence translation.