Malay Noisy Text Normalization in User-Generated Content: A Rule-Based and Levenshtein Distance Approach
Muhammad Fitri Shazwan Fadzely, Azilawati Azizan, Nursyahidah Alias, Nurkhairizan Khairudin, Rohana Ismail, Norshuhani bt Zamin · 2024
The rise of social media and informal communication has led to an increase in noisy text, characterized by abbreviations, spelling variants, incorrect spelling, mixed language, loan words, phonetic and slang. This poses a significant challenge for Natural Language Processing (NLP) tasks, which rely on clean and structured text. The widespread use of informal Malay language in user-generated content (UGC) on social media further exacerbates this issue. To address the normalization of such noisy Malay text, this study combines rule-based (RB) techniques with Levenshtein Distance (LD). The proposed approach specifically addresses three types of noisy text, including abbreviations, incorrect spelling and spelling variants. The development involved several phases: data acquisition and requirement analysis, pre-processing, noisy word identification, normalization process, and detokenization. The normalization process employs rules derived from linguistic patterns in Malay, combined with LD calculations to correct deviations from standard words. The effectiveness of this approach was evaluated using the Malay Chat-style-text Corpus (MCC) and recent UGC data from X application. The results demonstrate a significant improvement in text quality, with 51.1% of noisy words in the MCC and 64.6% in recent UGC being successfully normalized. This improvement is expected to significantly improve the performance of NLP applications in Malay, including sentiment analysis, various form of text analysis, and information retrieval applications in general.