Automatic Semantic Network Generation from Unstructured Documents – The Options
Barack Wanjawa, Lawrence Muchemi · 2018
There is plenty of information that exists in freeform or web-based text and whose use from a computing perspective is limited because of its unstructured nature. This data first needs to be structured as a knowledge base to facilitate tasks such query, search, re-use, sharing, keeping information current, question answering (Q&A), prediction, disambiguation and summarization. Structuring through topic models, database base systems and taxonomies have been tried but these tend to be domain specific and limited in scalability to the level of diverse unstructured data available. Knowledge representation to create knowledge bases can be done through linking data as used in applications such as LinkedIn, Facebook, BabelNet, Google Knowledge Graph, among others. This essentially forms a semantic structure also called a semantic network (SN). The automatic generation of these structures is a problem largely unresolved. This paper concentrates on development of a model that uses machine learning to automatically generate semantic networks from web data. SNs are an important component in the development of applications that support semantic search, Q&A, information retrieval among others. Automatic generation, also known as semantic network induction, reduces the complete reliance on manual and time-consuming semantic network authoring.