2 Challenges of Persian NLP: The Importance of Text Normalization

Katarzyna Marszałek-Kowalewska · 2023

Data is said to be the new oil. This applies to textual data as well. The amount of plain textual data on the Internet is abundant, and it is still growing. However, similar to oil, raw data is not valuable by itself. Its value and potential are availed through further accurate processing and analysis. And this is where Natural Language Processing (NLP) and its procedures come into the picture. Since NLP tasks require standardized and high-quality inputs for high efficiency and performance, raw textual data requires at least a basic level of cleaning and standardization. Therefore, text normalization - a process of transforming noisy data into improved, i. e., a standard representation - is often a prerequisite for a variety of NLP tasks, including information extraction, machine translation, or sentiment analysis. The Persian language, being the 5th content language of the Web, can be a great source of diverse information. However, several language-specific characteristics hinder the potential application of Persian data. The focus of this chapter lies in describing challenges that the Persian language poses for NLP and evaluating the impact of normalizing a large Persian corpus for one of the downstream NLP tasks - the discovery of multiword expressions (MWEs).

Read the paper · More papers on PaperTik