Towards Natural Natural Language Processing A Late Night Brainstorming
Petr Sojka · 2008
An essay about mimicking some aspects of language process- ing in our heads, using information fusion and competing patterns. Usual approach to Natural Language Processing separates language processing into word form, morphological, syntactic, semantic and pragmatic levels. Most often processing of these levels are independent, and result of one level is communicated to the other unnaturally disambiguated to cut off less probable (but often linguistically valid) intermediate results (e.g. sentence syntactical parse trees) just to simplify things. Even though ungrammatical sentences are often used for communication between people (English as the second language), they are banned by NLP software. Considerable effort is given to the balancing general purpose corpora to choosen only such text examples, aiming at handling only (syntactically) correct language parts. Given that, for purposes of handling non-polished texts, blogs or even speech, these data resources fail badly, as the tools are trained and fine-tuned to the different type of input than used when processing real (speech) data.