Ontology-Based Structured Cosine Similarity in Speech Document Summarization
Soe-Tsyr Daphne Yuan, Jerry Chih‐Yuan Sun · 2004
Development of algorithms for automated text categorization in massive text document sets is an important research area of data mining and knowledge discovery. Most of the text-clustering methods were grounded in the term-based measurement of distance or similarity, ignoring the structure of terms in documents. In this paper we present a novel method named Structured Cosine Similarity that furnishes document clustering with a new way of modeling on document summarization, considering the structure of terms in documents in order to improve the quality of speech document clustering. 1.