What do Asian and non-Asian scriptures have in common? An applied statistical machine learning inquiry
Preeti Sah, Ernest Fokoué · Mathematics for Applications · 2020
This paper presents a substantially detailed statistical machine learning approach to the analysis of several aspects of sacred texts from both the Asian and Biblical scriptural canons.The corpus herein considered consists of 4 Asian sacred scriptures, namely the Tao Te Ching, the teachings of the Buddha, the Yogasutras of Patanjali, and the Upanishads, and 4 non-Asian sacred texts essentially four books from the Bible, namely the Book of Proverbs, the Book of Wisdom, the Book of Ecclesiastes and the Book of Ecclesiasticus.Standard text mining tools are used, like the creation of Document Term Matrices (DTM) to pre-process raw English translations into word frequencies, and both unsupervised and supervised learning methods are used to answer some foundational questions featuring similarities and dissimilarities within each canon and interesting differences between all the canons considered.Despite the vast disparities between the translators of the original texts, our findings reveal sharp differences between Asian and non Asian scriptures regardless of whether clustering techniques or pattern recognition methods are used.We provide several compelling visualizations to help highlight our striking findings, chief of which are the persistent groupings of the scriptures based on geography. MSC (2010): primary 62F15; secondary 62F07.