Knowledge Graph for Deriving Insights on The Thai Government Dataset
K. Saratoon, A. Chutiporn, S. Nuttapong · 2023
Natural language processing (NLP) is mandatory in working with text. There are many tools and applications that are based on it. However, most of those tools are often operated in, or only, the English language. In recent years, there has been a continuing development for NLP tools to support other languages or creating a specific tool for certain language for simple NLP tasks, but for some of the advanced tasks, the advancement is still behind that of the English language, Thai language is also one of them. So, in this research, the capabilities of the currently existing Thai NLP tools are explored and evaluated, with the tasks of extracting text from the Thai government dataset (eMENSCR) and creating the knowledge graph from it to improve data interpretability and gain more insight from the data, by utilizing queries that are exclusive, or less complex to execute, when the data is stored in the graph database such as performing a path traversal or relationship counting on the data. Natural language processing's part of speech tagging and named entity tagging is used to perform entity and relation extraction after filtering the unneeded data fields. Then the extracted information will be formulated into the format of “triple”, which is in the form of (head, relation, tail). After the process of triple construction is finished, The triples are evaluated by Precision, recall, and F1 in order to measure the pipeline's performance and import to the Neo4j for query testing. The obtained results show that there is still room for improvement for both the tools and the methodology itself.