Performance of End-to-End vs Pipeline Spoken Language Understanding Models on Multilingual Synthetic Voice
Mohamed Lichouri, Khaled Lounnas, Rachida Djeradi, Amar Djéradi · 2022
This work conducts a comparative investigation of two architectures in the domain of Spoken Language Understanding (SLU), which were evaluated on a synthesized corpus of three languages: Modern Standard Arabic (MSA), French, and English. The first architecture employs a simple SLU system based on classical machine learning algorithms (E2E SLU), whereas the second architecture (Pipeline SLU) merges the textual output of a speech recognition system (ASR) with that of a textual classification system by transmitting it to a ”Natural Language Understanding” (NLU) model, allowing us to compare the predictions of the two systems. The obtained results were encouraging where we found that the Pipeline approach has given us better results than the E2E approach