How statistics and text mining can be applied to literary studies?
Mohammad Reza Mahmoudi, Ali Abbasalizadeh · Digital Scholarship in the Humanities · 2018
Abstract Statistics and data mining techniques provide exciting approaches for extracting knowledge from data. Recently, using statistics and data mining has sought to be exploited in many research fields. In this study, it was demonstrated that how statistics can be applied to literary studies. First, all the lines in Khaghani’s divan are classified and coded into three categories (mystical, non-mystical, and borderline). Then a set of chi-square goodness-of-fit tests are used to investigate and compare the frequency of different line’s categories for all lines and all odes, separately. Finally, the chi-square independence test (crosstabs) is employed to investigate the existence of trend in the lines.