A SEMANTIC APPROACH TO ANALYZE SCIENTIFIC PAPER ABSTRACTS
Ionut Cristian Paraschiv, Mihai Dascălu, Ștefan Trăușan-Matu, Philippe Dessus · eLearning and Software for Education · 2015
Each domain and its underlying communities evolve in time and each period is centered on specific topics that emerge from textual sources that characterise the domain. In this paper we propose a semantic analysis of the article abstracts extracted from the citation index Web of Science, from the category Education and Educational Research, taken between the years 2000-2004. Our analysis represents an extension of other researches performed on the same corpora that were focusing more on evaluating co-citations between the articles in order to compute their importance score (Jensen & Grauwin and Jensen, Grauwin, Lund & Jeong). Our approach presents a general perspective of the domain by performing semantic comparisons between article abstracts using natural language processing techniques such as Latent Semantic Analysis, Latent Dirichlet Allocation or semantic distances in lexicalized ontologies, i.e. WordNet. Moreover, graph visual representations are generated using Gephi in order to highlight the keywords of each paper and of the domain, making nevertheless possible the comparison of the articles marked as important to the ones extracted by Jensen, Grauwin and their colleagues. Overall, the articles' abstracts contain more relevant semantic information than co-citations and are less prone to errors while capturing the specificity of the domain. Also, in order to further argue the benefits of our approach, a debate about the advantages and disadvantages of these complementary methods is presented, together with potential refinements of the methods for classification that can be performed as future improvements. We hope that our research will give a better comprehension of the domain of Education and Educational Research.