Statistical Extraction and Visualization of Topics in the Qur'an Corpus

Maysum Panju · 2014

Unsupervised machine learning techniques are described and applied on the mildly preprocessed Arabic text of the Holy Qur’an, with promising results. Topic modelling based on nonnegative matrix factorization was used to successfully extract meaningful topics underlying the set of 6236 verses in the corpus. Data visualization using t-SNE dimensionality reduction correctly grouped verses of the Holy Qur’an into clusters based on theme and word usage. This accessible paper begins with an introductory view of machine learning, and includes motivating descriptions of the implemented techniques before presenting a summary of findings. A graphical display combining the results of topic modelling and data visualization demonstrates the consistency of the studied models.

Read the paper · More papers on PaperTik