Detection and elimination of similar Web pages based on text structure
Zheng Jia-heng · Journal of Computer Applications · 2007
Similar Web pages that search engine returns not only waste storage resources but also increase the burden on Web users.A dynamic method to detect similar Web pages was proposed.By this method,Texts of Web pages were expressed in the style of catalogue structure trees according to the features of similar Web pages and the features of Web pages themselves.A dynamic algorithm to extract features of texts and a layer fingerprint algorithm to calculate similar degree were implemented.The experimental results show that completely similar Web pages are detected accurately,and partly similar Web pages are detected exactly.