Contextual representation and learning for unsupervised knowledge discovery in texts
Patrick Perrin · Medical Entomology and Zoology · 1997
This dissertation studies the role of lexical contextual relations for the problem of unsupervised knowledge discovery in full natural language texts. Narrative texts have inherent structure dictated by language usage in generating them. We suggest that the relative distance of terms within a text gives sufficient information about its topic structure and its relevant content. Furthermore, we suggest that this structure can be used to discover implicit knowledge embedded in the text, and therefore serves as a good candidate to represent effectively the text content for knowledge elicitation tasks. We qualitatively demonstrate that a useful text structure and content can be systematically extracted by collocational lexical analysis without the need to encode any supplemental sources of knowledge. We present two algorithms that have been fully developed. The first is to systematically extract the most relevant facts in the texts and to label them by their overall theme, dictated by local contextual information in the texts. It exploits domain independent lexical frequencies and mutual information measures to find the relevant contextual units in the texts. The second is a learning algorithm for clausal discovery in a first-order clausal representation of the texts. It integrates an information-based interestingness measure to discover interesting causal textual patterns. We report results from experiments in a real-world textual database of psychiatric medical reports.