Web Information Retrieval Based on XML and N-level VSM
Ran Zhang · Computer Technology and Development · 2006
XML documents have well form,clear levels and analyses the structure easily.Convert HTML documents on Web into XML document,so can use DOM tree in Java to analyse the hierarchy of the documents.The documents can be divided into N level text paragraphs' content,which are represented by index term vectors.Using this method improve traditional vector space model,the N level VSM is achieved.And proved by the experiment,both recall and precision of the N level VSM are performing well than the traditional VSM.