Semi-Structured Web Information Extraction Process Based on XML
Ying Zhang · Journal of Changchun University of Science and Technology · 2007
As a development of extraction technology in B/S field,web information extraction is a universal application for the purpose of storing and retrieving increasing items or data.Taking Web pages as original data resources,Semi-Structured Web Information Extraction Process Based on XML described in this paper mainly researched implements of the recessive information extraction,generation of more structured and semantic-cleared storage forms with some open source products and technologies such as XML、JTidy data cleaning、Xquare-bridg and so on.