Automatic Extraction of Place Entities and Sentences Containing the Date and Number of Victims of Tropical Disease Incidence from the Web
Taufik Fuadi Abidin, Ridha Ferdhiana, Hajjul Kamil · Journal of Emerging Technologies in Web Intelligence · 2013
Many tropical disease incidences, such as leprosy, elephantiasis, malaria, dengue fever, are reported in online news portals. Online news portals are valuable data sources for creating a tropical disease repository if the information such as the location of the incidence, date of occurrence, and the number of victims can be automatically extracted from news articles. This paper describes approaches to extract that information from the Web. We introduce a rule-based algorithm to identify and extract the locations of the incidence and use Support Vector Machine (SVM) to determine the sentences containing the date of occurrence and the number of victims. Our experiments show that, the accuracy of the rule-based algorithm to identify the location entities is 99.8%, while the accuracy of the classifier to determine the sentences that contain one or more places of the incidence is 82%. The accuracy of SVM classifiers to classify the sentences that contain the date of occurrence and the number of victims are 96.41% and 93.38%, respectively.