Information Extraction from Arabic News
Hala Elsayed, Tarek Elghazaly · 2015
Information Extraction (IE) is concerned with finding of specific facts from collections of vast unstructured texts found in the web and in large documents. The Named Entity Recognition (NER) is a sub-problem of the Information Extraction (IE). The recent research in information extraction are growing and also there are interests in the Named Entity Recognition (NER) which helps in e xtracting the desired information from massive texts and hence extracting entities is an important task in the Natural Language Processing (NLP). T he Arabic Language needs to perform more researches in information extraction domain and hence we introduce this research. The experiment is concerned with extraction entities and entities relation extraction from the Arabic text. We used in our experiment text from Arabic news in Egyptian Arabic newswire. The paper introduced a method for extracting numerous unknowns using entity and entities relation from Arabic Corpus that is generated from Egyptian Arabic newswire to extract Information using the Named Entities and Entities Relation in Arabic language. The experiment contained nearly 625368 entries; the number of sentences was 36423 and the selecting sample was about 3400 sentences representing the crimes news. In the results we obtained some information that is considered a tool for a decision-maker in analyzing the text.