Intelligent Text Clustering Based on Semantics Similarity

Manhal Elias Polus, Thekra Abbas · 2020

Clustering text documents have become an increasingly important problem in recent years due to the availability of a huge amount of unstructured data in various forms, such as the web, social networks, and other information networks. It aims to organise and classify large document groups into smaller groups of meaning. This process is crucial because it is challenging to deal with a large amount and an increasing number of digital data. The documents are organised and classified to facilitate faster information retrieval (IR) and try to extract information with semantic knowledge. The process above enables the retrieve information, browsed and understood instead of clustering texts in the traditional way, i.e. compilation of data without descriptive concepts. As textual data has become a diverse set of vocabulary hence, there is an urgent need for text aggregation techniques based on semantic similarity is the primary solution to this problem as it is grouped into groups according to meaning rather than keywords. Several Papers that are using semantic similarity in various scopes. It has been reviewed in this research; some of them which are using similarity based on semantic using document cluster. For developing an effective and efficient clustering methodology to take care of the semantic structure of the text documents, and a compare of between them is presented according to (algorithms, tools, and assessment methods). Finally, extensive study and comparison of work are presented.

Read the paper · More papers on PaperTik