Research on Public Opinion Information Screening Model Based on Word2Vec and LDA Topic Clustering
Yaxian Qiu, Lei Qiao, Yidan An, Guo Jiaqi, Junyi Pang · 2024
In the current social environment of information explosion, real-time monitoring and screening of public opinion information is of great significance for governmental decision-making, corporate strategic planning and public security. This paper proposes an innovative public opinion information screening model based on Word2Vec and LDA topic clustering, aiming to achieve fast and accurate analysis of massive text data through advanced natural language processing technology. The construction of the model begins with the preprocessing of raw data, including word splitting and deactivation removal, to construct a high-quality corpus. Subsequently, topic clustering is performed using the LDA model to mine potential topics in the document collection and extract representative keywords for each topic cluster. Further, this paper employs a Word2Vec model to train word vectors, which converts words in text into numeric vector form, thus enabling quantification of relationships between words in vector space. In addition, the model includes a text filtering mechanism based on similarity ranking, which ensures the accuracy and relevance of the filtering results. In order to verify the effectiveness of the model, a corpus of topic-fitting and deviating texts is constructed, as well as the similarity between the texts screened by the model and the reference texts is calculated. The experimental results show that the model is able to significantly distinguish between texts that are relevant and irrelevant to the topic, proving its efficiency and accuracy in the screening of public opinion information. The model successfully identifies texts that are highly relevant to specific keywords from a large number of texts, and their relevance scores are much higher than those of randomly selected texts, thus verifying the practical application value of the model. This method not only improves the efficiency of public opinion information screening, but also verifies its accuracy through experiments, providing a new technical tool for the field of public opinion analysis.