Textual Similarity based on Proper Names

Nathalie Friburger, Denis Maurel · 2002

Proper names represent about 10% of English or French newspaper articles. Their quantity and informational quality is already used in different Information Extraction systems. Proper names have widely been studied in the MUC conferences designed to promote research in Information Extraction. We have created our own named entity extraction tool based on a linguistic description with automata. The extracted names are used in an information retrieval process: we want to cluster journalistic texts with a high precision level and to provide a description of the topic of the clusters. We verify the interest in the use of proper names in measures of similarities trying to improve the clustering of newspaper texts.

Read the paper · More papers on PaperTik