Text Generation & Classification in NLP
Kuldeep B. Vayadande, Dattatray Raghunath Kale, Jagannath E. Nalavade, Rajesh Kumar, Hanmant D. Magar · 2024
The initial stage in natural language processing is to break down the text into separate tokens. When the text corpus is huge, covering all words is inefficient regarding size of vocabulary. The effectiveness of a specific tokenization method varies on various factors, such as size of the dataset, the nature of the task, and the morphological complexity of the dataset. By comparing the algorithms, it can be concluded that no tokenization technique is the best choice. In this survey, various applications are being surveyed and the comparison of these various algorithms is done by estimating them on classification tasks like sentiment analysis. Question answering and translation applications use the available datasets. This survey paper also shows the tokenization based on the noisy text data and how various tokenization algorithm works on these data are being compared, and what is the average number of segmented subword accuracy being discussed. Basically, sentiment analysis studies the information in an expression and classifies them as positive, negative, or neutral. Input sentence is taken from the user. The survey shows that tokenization is the act of dividing a text into smaller parts like words, phrases, or sentences. The tokens are then used as the input to various NLP models. Tokenization helps to convert unstructured text data into structured data, which can be further processed and analyzed. For text classification, tokenization is used to convert the input text into a numerical representation. This numerical representation is then fed further, where the classifier predicts the label of the given text. Commonly used tokenization techniques for text classification include bag-of-words, n-grams, and word embeddings.