A Multiple Approximate String Matching Algorithm of Network Information Audit System
Peng Gao, Deyun Zhang, Qindong Sun, Zhai Ya-Hui, LU Wu-Chun · 2004
This paper shows a simple, efficient, and practical algorithm for locating all occurrences of a finite number of keywords in a char/Chinese character string allowing k chars inserting errors. The algorithm consists of constructing multiple finite state single-pattern matching machines from keywords and a state-driver applied to drive all finite state single-pattern matching machines, and then using the state-driver to process the text string in a single pass. Speed of the matching is independent of the amount of the inserting errors. Generally, the algorithms do not need to inspect every character of the string. They skip as many characters as possible by making full use of the information in matching failure and text window mechanism. This algorithm can be widely applied to network information auditing, database, information retrieval, and etc.