Proper Noun Extracting Algorithm for Arabic Language

Riyad Al–Shalabi, Ghassan Kanaan, Bashar Al-Sarayreh, Khalid Khanfar, Ali M. Al-Ghonmein, Hamed A. Talhouni, Salem Al-Azazmeh · 2011

Many of Natural Language Processing (NLP) techniques have been used in Information Retrieval, the results is not encouraging. Proper names are problematic for cross language information retrieval (CLIR), detecting and extracting proper noun in Arabic language is a primary key for improving the effectiveness of the system. The value of information in the text usually is determined by proper nouns of people, places, and organizations, to collect this information it should be detected first. The proper nouns in Arabic language do not start with capital letter as in many other languages such as English language so special treatment is required to find them in a text. Little research has been conducted in this area; most efforts have been based on a number of heuristic rules used to find proper nouns in the text. In this research we use a new technique to retrieve proper nouns from the Arabic text by using set of keywords and particular rules to represent the words that might form a proper noun and the relationships between them. To extract proper nouns from the retrieved document, we need some information about it and where it was found. First, we mark the phrases that might include proper nouns; second, we apply rules to find the proper noun and we use simple methods (stop wording and stemming) usually yield significant improvements. To test the system we have used 20 articles extracted from the Al-Raya newspaper published in Qatar and Alrai newspaper published in Jordan.

Read the paper · More papers on PaperTik