Learning a Domain Ontology from Hierarchically Structured Texts

Pavel Makagonov, Alejandro Ruiz Figueroa, Konstantin Sboychakov, Alexander F. Gelbukh · 2005

Any scientific or technical document is organized hierarchically: some sections of the text (such as the abstract or conclusions) summarize the contents of the main text; sections have titles describing their contents in general words; chapter titles describe the contents of a set of sections; book title describes the contexts of all chapters, etc. Moreover, whole collections of scientific documents are usually organized hierarchically: e.g., papers are organized in journals, conferences, etc., which in turn have their own titles. We exploit this hierarchical structure to learn a lexical ontology, in which subordination relationships roughly mirror those between the texts and titles in which these words occur: words occurring in more general titles subordinate the words occurring in the texts described by these titles. 1.

Read the paper · More papers on PaperTik