Statistical Keyword Matching using Automata
Seba Susan, Shashank Kumar, Rohen Agrawal, Kartik Yadav · International Journal of Applied Research on Information Technology and Computing · 2014
This paper proposes to statistically gauge the degree of matching of keywords in input strings using finite automata, in order to grade strings in the order of relevance with respect to the given keyword. The nonextensive entropy with the Gaussian information gain function proposed by Susan and Hanmandlu for the representation of regular patterns is used by us as the statistical measure. The improbable events falling in the ‘bell of the Gaussian information gain function’ are highlighted by this non-extensive entropy. The keywords which are the improbable events in a composite string are adequately represented by this statistical measure. The result is an improvised and more indicative grading than the combined pattern recognition tool of automata and reinforced learning used by several researchers.