Domain Topic and Hidden Deep Web Data Extracting

Liming Du, Abdulhamid Yahaya, Li Gui, Fengying Wang, Jie Dong · Proceedings of the 2018 3rd International Conference on Automation, Mechanical Control and Computational Engineering (AMCCE 2018) · 2018

This paper mainly studies the method of extracting web data entities based on domain.Through the analysis of real estate industry websites, a topic-oriented topic extracting model is proposed, and the corresponding search strategy is given.In addition, for the case of depth information, a sorting-based classification extraction algorithm is designed for numerical data.Finally, an experimental example is given to verify the effectiveness of the algorithm. Domain Topic Extracting Model and Searching Strategy Domain Topic Extracting Model.The domain topic extracting model of this article has been improved on the basis of the generic crawler model, and the flowchart of the extracting model used in this paper is shown in the Fig .1.Compared with the general crawler model, the domain topic model has two more modules: the page topic relevance calculation module and the candidate URL priority calculation module.The page topic relevance calculation module may filter the saved pages according to the relevance of the pages and the topics.If the relevance of the page to the topic is higher than the set threshold, the candidate URL of the page is extracted and input into the candidate URL priority calculation module, and the calculation rules are as follows: If the candidate URL is relatively related to the topic, it is inserted.To the front of the queue, the opposite is inserted into the back of the queue or is discarded.

Read the paper · More papers on PaperTik