Data-driven approach for ontology learning
Isidra Ocampo-Guzman, Iván López-Arévalo, Victor J. Sosa‐Sosa · 2009
In this paper, an approach to construct ontologies based on a text corpus is described. Using latent Dirichlet allocation the topics that describe the documents contained in the corpus are identified. Each topic is formed by a set of terms whose semantic relatednesses are determined applying the distributional hypothesis, which considers as similar terms those that share the similar linguistic context. This context is described by the verbs they share. The concept described by each topic's terms is modeled through a taxonomy that describes the relation between them.