Focused crawler for the acquisition of health articles
Amalia Amalia, Dani Gunawan, Atras Najwan, Fathia Meirina · 2016
The health intervention by using technology can be the alternative to the doctor, especially for common health problem. To support the technology, we need health knowledge base as the foundation. The artificial intelligence and hardware development nowadays support this requirement. The big picture of our research is building the application that can utilize the health knowledge base to provide health intervention. As the first step, we collect the articles related to health. To realize it, we build the focused crawler that implements multithreaded programming, Larger-Sites-First algorithm and also Naïve Bayes classifier. We find that the articles acquisition is going to saturate along with the increment of threads. Furthermore, the implementation of Larger-Sites-First algorithm do increase the number of crawled articles, but it is not significant. In addition, Naïve Bayes recognizes ≥ 90 percent articles in perfect condition for both health and non-health category. However, the performance goes down when recognizing the non-health articles which contain health keywords.