Document clustering based on similarity of subjects using integrated subject graph
Masao Nakada, Yuko Osana · International conference on Artificial intelligence and applications · 2006
In this research, we propose an integrated graph which expresses the of the document. The proposed integrated graph is based on the graph-based text representation model which is called subject In the graph, a node represents a term in the text, and an edge denotes a relation between linked terms. As the conventional text representation models, the graph models such as the graph and the KeyGraph have been proposed, and most of them assume that one document has one subject. However, the document often has not only one but also plural subjects. In this research, we assume that each unit of the document such as a paragraph has one subject, and each unit is translated into a graph. Then, they are integrated into an integrated graph. In this research, we apply the proposed integrated graph to the document clustering and realize the document clustering based on the similarity of the subjects. We carried out a series of computer experiments and confirmed the effectiveness of the proposed integrated graph.