Email based categorizationon surname using suffixtree algorithm

Reshma Suresh, P Abirami, Archanakumari Sharma, T Kavitha · 2017 International Conference on Computation of Power, Energy Information and Commuincation (ICCPEIC) · 2017

Data mining is a process of examining data from various perspectives. Here, we use a dataset which contains collection of email-id having forename and surname within it. The email-id are combined together as a record with spelling errors. Absence of identification or space may be present among forename and surname that forms a endless string. Here, we have taken the problem of identifying the surname applying various algorithm specially the suffix tree algorithm. At first, we proposed a correction technique for correcting the spelling errors of the forename and also the surname using Edit-distance algorithm. And next, we proposed Phonetic search, which is used as a searching strategy based on pronunciation of words and then we have used a suffix tree algorithm, it is a compressed tree which contains all the suffixes of the text as their keys. Suffix trees allows fast execution of different string operations a, n clustering algorithm, hierarchical clustering technique. We experimented on Indian email id dataset and modified Jaro-Winkler technique for correction and performed Suffix tree algorithm for faster computation and results are summarized in this paper.

Read the paper · More papers on PaperTik