Application of a Mining Algorithm to Finding Frequent Patterns in a Text Corpus: A Case Study of the Arabic

Imran Ali, Klong Luang · 2012

Information repositories containing text data of different languages are abundant on the World Wide Web. Digital corpora of sacred text of Islam related to Quran containing Arabic language are also publicly available. The availability of these corpora and intelligent application to analyze them are vital to better comprehend the religious text of Islam. In this paper I propose a method of representing the Quranic text corpus as a graph, and apply a frequent sub-path mining algorithm on it to generate frequent patterns. I have explained how the resulting frequent patterns can be used for subjective indexing and clustering similar verses of Quran.

Read the paper · More papers on PaperTik