Topic Detection in Domain Text Dataset: An FCA-Based Approach
Fugang Wang, Nianbin Wang, Shaobin Cai, Hongbin Wang · Journal of Circuits Systems and Computers · 2024
Text data contain a large amount of useful information, and various domain texts obtained after text data classification are the main body of text data research. Domain text data share a common domain background, so the topics in a domain have more overlap and finer differences. The concept in formal concept analysis (FCA) is a bi-aggregation structure of objects and attributes, which is regarded as the unit of human thought. The generation-specialization relationship between concepts, which is a both bidirectional containing and independent relationship, presents a hierarchical structure, which is very suitable for describing the significant or subtle similarity relationship between real categories. Due to the large scale of concept lattices, it is a challenge for FCA theory to deal with domain text data with the characteristics of large scale, no label and long text. This paper proposes a natural clustering method based on FCA, which can effectively discover, in the domain text data, all the topics with the granularity that users are interested in and can flexibly set the interested granularity of the topics. The feature extraction algorithm for domain texts is adopted, which not only effectively reduces the dimension of attributes and highlights the topic characteristics of documents (objects), but also eliminates the influence of noise in data. It makes the structure of the formal concept more concise, the semantics of that more clear, and the topic reflected by that more explicit and interpretable. The comparison with the LDA model on real data shows that our method is an effective, better structured and more interpretable topic detection method.