HIGH ACCURACY POSTAL ADDRESS EXTRACTION FROM WEB PAGES
Zheyuan Yu · 2007
Automatically extracting geographic information fromWeb pages, and analyzing such information can greatly benefit data mining and information retrieval systems. One most obvious but valuable source of geographic information is postal address. In this thesis we present three methods for postal address extraction - rule-based, machine-learning and hybrid. Our machine-learning based system combines different sources of weak evidence and uses the word n-gram model as the underlying rep- resentation of web pages. The extracting process is fast and accurate despite not using dictionary look-up. It outperforms both rule-based systems that we built, with precision of 94.3% and recall of 71.6%. We also proposed a hybrid method which combines the rule-based method and the machine-learning approach. The accuracy of the system is further improved with precision of 95.2% and recall of 81.1%.