Topic Modelling in Computer Security Discourse: a Case Study of Whitepaper Publications and News Feeds

Ekaterina Isaeva · Вестник Пермского университета. Российская и зарубежная филология · 2022

Up-to-date information plays a crucial role in modern linguistic research. For this reason,computational linguistic methods, including those aided with analytical and machine-learning tools, are attracting growing attention. Some of their applications in cognitive-discursive linguistics are keyword extraction, topic modelling, and content analysis. Text-mining tools facilitate time-consuming linguistic work andadd to the results’ reliability and greater statistical precision by processing a significantly larger data volume.Most studies, however, have overlooked interference of socially significant but context-irrelevant (e.g. political) information into a specialized discourse by focusing mainly on one data format. The current study,aimed at topic modelling, has been carried out on the computer security discourse. We have implemented theproject on the KNIME analytical platform. The model enables comparison between topics extracted frompublished articles and date-specific RSS news feeds. The study provides important insights into infodemiology and political incidental news exposure occurring in computer-security-oriented RSS feeds on theKaspersky website but untraceable in the papers published on the same website in a PDF format. The resultsreported here provide further evidence for the need to consider the hypercontext of professional communication and employ real-time data in solving similar problems within cognitive-discursive linguistics.Our contribution to the development of cognitive-discursive linguistics is the method for comparingtopics within one discourse, taking into account near-real-time data. For computational linguistics, the significance of our work lies in describing a new application of the topic extraction workflow freely available onthe KNIME hub.

Read the paper · More papers on PaperTik