Building the Indonesian NE Dataset Using Wikipedia and DBpedia with Entities Expansion Method on DBpedia

Haji Dito Murya Alfarohmi, Moch Arif Bijaksana · 2018

In Indonesian, the NER (Named Entity Recognition)system still needs a lot of improvement. Though NER is the main component in IE (Information Extraction)which is used by other advanced components. To create a reliable Indonesian NER system using a machine learning approach, large dataset is needed. If the dataset is constructed by tagging it manually, the size of the dataset generated is very small. Therefore, a system was created to build Indonesian NE (Named Entities)dataset which were tagged automatically using Wikipedia data as a source of corpus and DBpedia as NE labeling reference with the Entities Expansion method to expand DBpedia NE labeling reference. Currently, the existing system cannot detect name that contain words beginning with lowercase letter on automatic tagging, the existing system have not tried adding person entity gazetteers, and the DBpedia Entities Expansion method rules can still be modified to produce better NE labeling reference quality. In this study a system was built to overcome these shortcomings. Evaluation showed that the best Indonesian NE dataset was built in this study produced Fl-score of 54.93 %, 3.32 % higher than the result of previous studies 51.61 %. This best dataset was built by adding a detection method on automatic tagging, that using the DBpedia Entities Expansion modification rules in this study, but without adding person entity gazetteers.

Read the paper · More papers on PaperTik