Benchmark Text Preprocessing Techniques in Natural Language Processing
Aditi Tyagi, V K Jain, Vivek Kumar · 2024
Text preprocessing is a key step in Natural Language Processing (NLP) that deals with the cleaning, tokenization and structure of text before building models. A comparison of the recent advancements conducted based on their algorithmic concepts and efficiency of computations. Text preprocessing techniques like tokenization, normalization, stemming, and lemmatization are used in a range of NLP applications. In this study, we synthesize results from recent literature identifying strengths and weaknesses of the different approaches across languages and domains. We used five preprocessing algorithms for benchmarking, based on intrinsic, extrinsic, efficiency, interpretability, and human evaluation. We found that SpaCy Tokenizer excelled in lexical diversity, NLTK Tokenizer in readability, and BERT Tokenizer in information metrics.