Toward a taxonomy of concepts using web documents structure

Rim Zarrad, Narjes Doggaz, Ezzeddine Zagrouba · 2012

Due to the rise of the Web and the need to have structured knowledge, an interesting line for research is the formalization of ontologies and the creation of conceptual taxonomies from Web documents. The traditional methods for ontology learning and especially those extracting domain concepts from a textual corpus often privilege the analysis of the text itself, whether they are based on a statistical or linguistic approach. In this paper, we propose an approach which differs from the traditional ones since it uses information on the document structure to extract relevant information. Our approach studies each material form in the text in order to extract the most relevant concepts constituting the ontology related to a given field. The concepts are obtained by analyzing the occurrences of the candidate terms in the titles and in the links belonging to the documents and by considering the used styles.

Read the paper · More papers on PaperTik