Research on Web Information Extraction Model Based on XML and DOM Technologies
Wu Deng · Journal of Dalian Jiaotong University · 2013
XML technology is applied in search engine,and a web information extraction model based on XML and DOM technology is proposed.The stages of data acquisition,web age optimization,extraction rule generation and information extraction are analyzed in detail.The technologies of webpage reptile,NekoHTML,Xerces-J,JTree,Xpath and XSLT are applied in Web information extraction.Finally,semi-automation method of Web information extraction is realized.