Traditional Natural Language Processing

Field Cady · 2024

This chapter discusses techniques for natural language process, especially the ones that pre-date large language models. It starts with bag-of-words – the central concept for turning text into a numerical vector – and covers increasingly sophisticated ways of using linguistic concepts (lemmatization, synsets, etc.) to improve the vectorization process. Classic problems like sentiment analysis and topic modeling, which you can do once the data is vectorized, are covered along with more advanced topics like syntax trees and ontologies.

Read the paper · More papers on PaperTik