Spoken language identification: An overview of past and present research trends

Douglas D. O’Shaughnessy · Speech Communication · 2024

• Analysis of speech signals for automatic estimation of the language spoken. • Automatic speech recognition, speaker verification, and language identification are compared. • What distinguishes different spoken languages is discussed. • Utility of methods is noted in terms of performance, using accuracy, complexity, and cost as measures. • Approaches include: phonotactics, use of intonation, mel-frequency cepstral coefficients, neural networks. • Major components of neural systems (CNN, RNN, Transformer) are discussed. Identification of the language used in spoken utterances is useful for multiple applications, e.g., assist in directing or automating telephone calls, or selecting which language-specific speech recognizer to use. This paper reviews modern methods of automatic language identification. It examines what information in speech helps to distinguish among languages, and extends these ideas to dialect estimation as well. As approaches to recognize languages often share much in common with both automatic speech recognition and speaker verification, these three processes are compared. Many methods are drawn from pattern recognition research in other areas, such as image and text recognition. This paper notes how speech is different from most other signals to recognize, and how language identification differs from other speech applications. While it is mainly addressed to readers who are not experts in speech processing (as detailed algorithms, readily found in the cited literature, are omitted here), the presentation covers a wide discussion useful to experts too.

Read the paper · More papers on PaperTik