Generating Interpretable Features for Context-Aware Document Clustering: A Cybersecurity Case Study

Luis Felipe González Gutiérrez, Akbar Siami Namin · 2022 IEEE International Conference on Big Data (Big Data) · 2022

This paper extends the applications of ContextMiner, a framework that we proposed recently aiming to extract interpretable contextual features to conceptualize knowledge in security texts. We show that contextual features are suitable for building a document-level representation that can be used further in a downstream machine learning and natural language processing task: document clustering. Our results show that the intrinsic readability of contextual features is usable for analyzing the document clusters obtained in our experiments. Such analysis is performed through a density-based feature selection, cluster visualization, and statistical methods in order to unveil which contextual features better characterize each document cluster. Our findings suggest that statistical techniques alongside feature analysis can be utilized to discover meaningful commonalities among documents in a particular cluster without the need of querying each document manually.

Read the paper · More papers on PaperTik