A novel web page text information extraction method

Chongjun Wang, Wei Peng · 2019 IEEE 3rd Information Technology, Networking, Electronic and Automation Control Conference (ITNEC) · 2019

The text information of today's mainstream web pages generally has multiple features, and can be divided into a single text page and a multi-text page. In order to extract the text information of the web page, the position of the text information can be accurately located by using the multiple features of the text and the rules of the web page design. According to the above characteristics, this paper proposed a method for extracting web page text information based on multi-feature fusion. Experiments based on a large amount of data showed that the method has universality and high accuracy for the text information extraction of single text and multi-text web pages, and is very suitable for web pages with various styles.

Read the paper · More papers on PaperTik