A Novel PSS Stemmer for String Similarity Joins
P. Selvaramalakshmi, S. Hari Ganesh, J. James Manoharan · 2017
String similarity join plays a vital role in the integration and cleaning of data by identifying the similar pair of strings from one or more sources. The focus on string similarity join has been increased in the current scenario as it emphasizes on the effective management of data for the accurate analysis of hidden truth. String similarity joins always comes under the scope of stemming techniques as they concentrate on the explication of the syntactic base of a given string which would be helpful for the mapping of data from multisets. The string similarity join algorithms proposed in the existing literature have not been thoroughly experimented to prove its real time usability. Hence, the objective of the paper is to introduce a novel stemming algorithm for identifying prefix and suffixes of the given word to be used for the similar strings from one or more sources and to prove the real time usability of the proposed work.