Searching Proper Names in Databases.

Ulrich Pfeifer, Thomas Poersch, Norbert Fuhr · 1995

Identifying names --- e.g., author names or company names --- is still an open problem. In this paper we review known similarity measures. These measures deal with phonetic similarity, typing errors and plain string similarity. We show experimentally that all three approaches lead to significant better retrieval quality than plain identity. Furthermore, we demonstrate that combinations of different similarity measures perform even better than any single technique. 1 Introduction In our view modern information retrieval systems are characterized by being capable of dealing with the two features vagueness of queries and uncertainty of knowledge. In this paper, we focus on searching proper nouns. Here the vagueness stems from limited knowledge a user has about his information need. If he is e.g. searching for papers of a certain author in a bibliographic database, he will not be successful if he misspells the authors name in his query. The other item, uncertainty of knowledge, refers to ...

Read the paper · More papers on PaperTik