Identifying the Semantic Relations on Unstructured Data

Chien D.C. Ta, Tuoi Phan Thi · International Journal of Information Sciences and Techniques · 2014

Ontologisms have been applied to many applications in recent years, especially on Sematic Web, Information Retrieval, Information Extraction, and Question and Answer.The purpose of domain-specific ontology is to get rid of conceptual and terminological confusion.It accomplishes this by specifying a set of generic concepts that characterizes the domain as well as their definitions and interrelationships.This paper will describe some algorithms for identifying semantic relations and constructing an Information Technology Ontology, while extracting the concepts and objects from different sources.The Ontology is constructed based on three main resources: ACM, Wikipedia and unstructured files from ACM Digital Library.Our algorithms are combined of Natural Language Processing and Machine Learning.We use Natural Language Processing tools, such as OpenNLP, Stanford Lexical Dependency Parser in order to explore sentences.We then extract these sentences based on English pattern in order to build training set.We use a random sample among 245 categories of ACM to evaluate our results.Results generated show that our system yields superior performance.

Read the paper · More papers on PaperTik