A Hybrid Approach for Measuring Similarity between Government Documents of China
Zeyuan Li, Jie He, Dagang Chen, Xin Fang, Yajun Song, Zesong Li · Proceedings of the 2018 2nd International Conference on Computer Science and Artificial Intelligence · 2018
In China, the government publishes hundreds of thousands of government documents every year. The civil servants in China are struggling to find relevant government documents while doing their office works, such as writing documents, analyzing government policy, explaining the policy to the public. Furthermore, the public also finds it difficult to find the exact government documents since the most popular search engines in China such as Baidu, Sogou are not specialized in the field of government document searching and indexing. Determining the similarity between documents is critical to applications such as search and recommendation. Currently, most kinds of literatures focus on semantic similarity between words and paragraph fragments. As for government documents in China, they are written under standards so that they can be considered as semi-structured data after data cleansing. In this paper, we propose a hybrid approach for measuring the document-level similarity between government documents of China. We represent government documents as the publisher of the document, the document domain, the document type, the publishing time and other contents. By calculating the similarity between the elements, the government documents similarity is formed by the weighted sum of the similarity between the elements. Experiment results show that the proposed hybrid method outperforms the classic methods like LDA and doc2vec.