Automatically Extracting Structured Data from Web Pages of Similar Structure or Layout

Zhao Jing, Wang Qiao-wen, Guan Ma-zhou, Shan Chuan-jia · Journal of Anhui Science and Technology University · 2010

Database-driven web sites generate HTML pages in similar structure or layout.Traditional web information extraction methods often neglect or fail to use this similarity directly,so their efficiency and precision are generally poor.In order to exact structured data from structure-alike web pages,we presented a new model,Tag-Tree similarity model,which extends the conceptions and algebra of traditional embedded set,and a template method based on Tag-Tree for extraction in this paper.The result of our experiments shows that it is quite precise and effective.

Read the paper · More papers on PaperTik