Distributed Representations of HTML Page

Peihuang Wu, Jiakun Zhao · 2022 IEEE 2nd International Conference on Power, Electronics and Computer Applications (ICPECA) · 2022

The most common ways to represent documents as fixed length vectors are the bag-of-words model and the Doc2Vec model. The bag-of-words model mainly has two disadvantages. One is that it ignores the sequence and frequency of each word in the document, which makes the vectors lose the original semantic information; another is that the vectors of the bag-of-words model are high-dimensional sparse. Using these vectors as the input of the algorithm will greatly reduce the convergence speed of the algorithm. The Doc2Vec model overcomes the shortcomings of the bag-of-words model, makes the mapped document vectors have real semantic features, and ensures the low-dimensionality and denseness of the vectors. However, the Doc2Vec model is not very suitable for web pages on the Internet. The main reason is that web pages are usually semi-structured HTML format, containing many tags unrelated to specific content, and the meanings represented by each tag are very different. Therefore, based on the Doc2Vec model, this paper proposes a method of mapping HTML pages to fixed-length feature vectors — Html2Vec. The experimental results show that for HTML pages, using the Html2Vec method proposed in this paper to train the embedding vectors is better than using the Doc2Vec model or the bag-of-words model. The trained embedding vector can be applied to web page classification and other fields.

Read the paper · More papers on PaperTik