Research on Web Information Extraction Model Based on XML and DOM Technologies

Wu Deng · Journal of Dalian Jiaotong University · 2013

XML technology is applied in search engine,and a web information extraction model based on XML and DOM technology is proposed.The stages of data acquisition,web age optimization,extraction rule generation and information extraction are analyzed in detail.The technologies of webpage reptile,NekoHTML,Xerces-J,JTree,Xpath and XSLT are applied in Web information extraction.Finally,semi-automation method of Web information extraction is realized.

Read the paper · More papers on PaperTik