A History of Language Modeling

Justin Hutchens · 2024

Nearly all Natural language processing (NLP) systems to date could be sorted into one of two categories: rule-based NLP and statistical NLP. For rule-based models the creator gives the system the rules of interaction, and in statistical models the creator gives the system a framework to create its own rules of interaction. One hugely important feature that was employed to improve both the usefulness and the “humanness” of rule-based NLP was the use of pattern matching techniques. Another effective tool that developers have used to streamline the operations of rule-based NLP systems is input preprocessing. In computer science, exception handling is a common mechanism used to ensure the reliability of a system. Many of the earliest useful statistical language model systems were based on the n-gram approach to analyze text data.

Read the paper · More papers on PaperTik