Research on Topic Analysis Based on Zhihu Web Crawler—A Case Study of “Anxiety, Depression, and Confusion”
东山 杜 · Computer Science and Application · 2025
在信息时代,知乎作为重要的知识问答社区,积累了海量数据,为多学科研究提供丰富资源。本研究通过构建一个高效稳定的知乎数据爬虫系统,从知乎平台针对“抑郁、焦虑、迷茫”三个话题抽取了701条问题数据,去除了专栏文章和电子书后,保留了575个有效问题和220,184条有效回答,再对每个回答的内容进行解析,提取问题信息,包括问题ID、话题、问题描述和回答数等信息,在此基础上采取关键词分析、情感分析、影响力分析、性别差异分析方法并对结果进行了讨论。本研究的结果不仅对理解网络媒体上心理健康话题的讨论和影响具有重要意义,还为相关领域的研究提供了实证数据和方法论支持。In the information age, as an important knowledge Q&A community, Zhihu has accumulated massive data, providing rich resources for multidisciplinary research. In this study, an efficient and stable Zhihu data crawler system was constructed to extract 701 question data from the Zhihu platform for the three topics of “depression, anxiety, and confusion”. After removing column articles and e-books, 575 valid questions and 220,184 valid answers were retained. Then, the content of each answer was analyzed to extract question information, including question ID, topic, question description, number of answers, etc. On this basis, keyword analysis, emotional analysis, influence analysis, and gender difference analysis methods were adopted, and the results were discussed. The findings of this study not only hold significant implications for understanding the discussions and impacts of mental health topics on online media but also provide empirical data and methodological support for research in related fields.