Naïve bayes based language-specific web crawling
Ekkasit Srisukha, Supakpong Jinarat, Choochart Haruechaiyasak, Arnon Rungsawang · 2008
In this paper, we propose a Thai language specific Web crawling as a method of selectively seek out Web pages written in Thai. The strategy is to follow a URL with the highest probability of leading to Thai Web pages. The probability score is calculated from the example set of Web pages using simple Naive Bayes approach. In addition, we also use a heuristic based method to bias the probable URLs whose hosts have previously provided Thai Web pages. An experiment illustrated that the proposed method produces a high harvest rate and achieves a better coverage than the others.