Information Extraction based on Named Entity for Tourism Corpus

Chantana Chantrapornchai, Aphisit Tunsakul · 2019

Tourism information is scattered around nowadays. To search for the information, it is usually time consuming to browse through the results from search engine, select and view the details of each accommodation. In this paper, we present a methodology to extract particular information from full text returned from the search engine to facilitate the users. The approach is based on name entity recognition (NER). The main steps are 1) building training data and 2) building the model. The key task is the building training data: First, the tourism data are gathered and the vocabularies are built. Several minor steps include sentence extraction, relation and name entity extraction for tagging purpose. Then, the recognition model of a given entity type can be built. From the experiments, given hotel description, the model can extract the desired entity,i.e, name, location, facility as well as relation type. The extracted data can further be stored as a structured information, e.g., in the ontology format, for future querying and inference. The model for automatic named entity identification, based on machine learning, yields the error ranging 5%-25%.

Read the paper · More papers on PaperTik