Lexical Complexity Detection and Simplification in Amharic Text Using Machine Learning Approach
Gebregziabihier Nigusie, Tesfa Tegegne Asfaw · 2022
Text complexity is the level of difficulty of the document for understanding by the target readers. One common type of this text complexity is Lexical complexity which can cause comprehensibility and understandability problems for second language learners, and children, furthermore it is challenging for NLP applications. To reduce this text complexity for low resourced and morphologically reach language Amharic, we have designed a complexity detection and lexical simplification model using a machine learning approach. For the tasks, we have developed three subsequent models. The first model is used to classify text complexity, trained using 19k sentences. The second model is developed for detecting specific complex terms. The model is developed using 1002 unique complex terms. Lastly, we have trained word2vec (CBOW) model using 57.6k sentences (which contains 9756 unique tokens) for simplification generation and ranking. The experimental result of the classification model scores an accuracy of 88%(LSTM), 88%(BiLSTM), and 91%(BERT). The simplification generation for the identified complex term using cosine similarity results 92% for top-ranked and 61% for lowest-ranked simplest equivalents from five top-generated words.