SPlice: Automated Testing for Speech Translation via Syntactic Analysis

Z. Li, Ji Qi, Pin Ji, Jia Liu, Yang Feng · 2025

With the advancement of Deep Learning, the performance of speech translation systems has made remarkable progress. However, similar to traditional software, speech translation systems can still suffer from software defects that can lead to incorrect translations with potentially serious consequences. These systems are also vulnerable to real-world environmental interference, making their behavior unpredictable. The black-box nature of deep neural networks renders traditional testing methods ineffective, while the lack of diverse test cases and the challenge of constructing test oracles further hinder the implementation of their testing. To address this, we introduce syntactic structure invariance, a linguistically inspired concept that captures the structural containment between a pair of derivationally related sentences and their corresponding translations. Based on this concept, we propose a novel speech translation testing method, SPlice. SPlice simulates environmental disturbances on a seed speech, disassembles it into a template and speech blocks, and then generates multiple derivational speech pairs by inserting the blocks back into the template. SPlice detects translation errors by checking whether the syntactic structure invariance relation is violated in the translation results corresponding to the speech pairs. To validate SPlice, we experiment with three industrial speech translation systems: Google Translate, Youdao Translator, and Iflytek Translator. With 600 speeches crawled from the BBC as seed tests, SPlice detects 1,640, 1,101, and 1,305 translation errors with around 90.7% precision. The experimental results show that SPlice can effectively detect errors in the speech translation results with high precision, providing valuable information for developers to improve system performance.

Read the paper · More papers on PaperTik