Integrating NLP with Corpus Linguistics and Vice Versa

Sultan Almujaiwel · 2018

This paper is a call to bring Natural Language Processing (NLP) and Corpus Linguistics (CL) together in order to promote more effective linguistic research. Arabic NLP has recently come up with a significant number of techniques and methods that help to design, model and retrieve Arabic data for the purposes of text mining and language modelling. CL has itself recently developed beyond its established focus on frequencies, collocations, n-grams, distributions and statistical CL. This paper shows the urgency of developing a new path that integrates these two fields. The reason for such an integration between these two disciplines is that the former focuses on the idea that if there are no data there is nothing, while the latter focuses on the idea that if there is no data analysis no sense can be made of the linguistic data itself. The example that is given in this paper to exemplify the importance of such an integration is arTenTen1. This corpus has been manipulated and designed by a technique adapted from NLP but needs further computational process models to ensure the avoidance of dirty data.

Read the paper · More papers on PaperTik