Working with Raw Text
Taweh Beysolow · Apress eBooks · 2018
Those who approach NLP with the intention of applying deep learning are most likely immediately confronted with a simple question: How does a machine learning algorithm learn to interpret text data? Similar to the situations in which a feature set may have a categorical feature, we must perform some preprocessing. While the preprocessing we perform in NLP often is more involved than simply converting a categorical feature using label encoding, the principle is the same. We need to find a way to represent individual observations of texts as a row, and encode a static number of features, represented as columns, across all of these observations. As such, feature extraction becomes the most important aspect of text preprocessing.