An Automatic Discovery Framework of Frequent Topics, Association Rules, Clusters, and Sequential Patterns in Social Media with Chinese Word Segmentation
Vance Chiang-Chi Liao · 2023
More data are created in social media representing situations in the real world. Four hundred thousand posts on Facebook were collected and added to the mysql database to select the content to store them in comma-separated value (CSV) files. This work parted the data in order to use Chinese word segmentation system. Then, the data were tagged for the contents of posts in the Chinese word segmentation system. C# was used to parse the tagged result and select common nous (Na) and proper nouns (Nb) to count and the frequent pattern growth (FP-tree and FP-growth) algorithm was used to discover the frequent patterns. The result shows the top-k objects of the maximal counts and the association rules of the objects. The K-medoids algorithm was used to cluster the objects, and the efficient PrefixSpan (Prefix-Projected Growth) algorithm was used to discover their sequential patterns. This work is the first step for knowledge discovery and data mining of frequent topics in social media and these algorithms are original types. Various novel algorithms of topics can be proposed for mining new variations of association rules, clusters, and sequential patterns within this framework.