Improved Multilingual text Identification using Embedding Visualization and Deep Learning techniques

Arinjay Wyawhare, Aadiptya Basuli, Saikat Das, Rounak Jana, Purnendu Dutta Jha, Pravin Kumar Samanta · 2024

Classification of multilingual text is a difficult subject that has attracted a lot of interest lately. This study compares various methods for language recognition and classification that make use of embedding visualization and deep learning techniques. A dataset with text and language columns in 17 distinct languages is used in this work. Sentence Transformer is utilized for embedding, and built-in modules like langdetect, langId, and fasttext are utilized for language identification. The embeddings are visualized after dimensionality reduction with t-SNE, while classification is carried out with CNN, LSTM, and MLP models. Based on the results, the FastText MLP model has the highest accuracy of 0.998, precision of 0.998, recall 0.9980, and F -1 score 0.9985. However, the Sentence Transformer multi-layer perceptron model has an F -1 score of 0.956, recall of 0.957, precision of 0.958, and an accuracy rate of 0.957. Since FastText embeddings were trained on a large multilingual corpus, they are clearly clustered in two dimensions; thus the number of dimensions of the embeddings impacted language settings. It can be argued that for language classification, the FastText model works best with 16-dimensional embeddings.

Read the paper · More papers on PaperTik