Ontological knowledge extraction from natural language text
Syed Tauhid Zuhori, Md. Asif Zaman, Firoz Mahmud · 2017
Ontology (Onto=Being and Logy=Knowledge, therefore the Knowledge of Being) has a significant impact on the study of natural language processing. By providing a formal representation of knowledge it ensures proper understanding of a particular domain. A comprehensive description to build a vocabulary on the given domain can never be feasible without ontologies; in other word conceptualizations. There have been a number of works in the recent times to perceive the ontological knowledge to build a strong vocabulary. Most of the existing ontology construction tools support construction of ontological relations (e.g., taxonomy, equivalence, etc.). But the main problem is that they do not support construction of domain relations, non-taxonomic conceptual relationships (e.g., causes, caused by, treat, treated by, has-member, contain, material-of, operated-by, controls, etc.) which are basically found in the text sources. The first notable work on this field is a Named Entity Recognition system developed by Stanford University. Stanford NER (also known as CRFClassifier) is a Java implementation of a Named Entity Recognizer. It can successfully recognize at most seven classes. In this research work we have utilized this NER, and proposed an algorithm which includes POS tagging, lemmatization, parsing and pronoun and co-reference resolution. We can squeeze out 22 classes from these seven primary classes. We have compared our work with an existing system named TextOntoEx. On the basis of performance matrices we have analyzed both of the works on a same natural language text. Result shows that our proposed system can find out more classes than TextOntoEx.