Review of: "RelTopic: A Graph-Based Semantic Relatedness Measure in Topic Ontologies and Its Applicability for Topic Labeling of Old Press Articles"
Silvio Peroni · 2021
relatedness- measure-topic-ontologies-and-its-applicability-0 Review I want to thank the authors for having addressed the main part of my comments [1] .They have extended several parts of their work both in the text and in the experiments done.As a result, the article has been improved a lot.However, some issues (listed below) should be addressed to have the article accepted in the Semantic Web journal.Inter-rater agreement By reading Section 8.2, it is unclear how the authors have computed the inter-rater agreement among the three annotators and the results obtained.In particular, it is not clear how the various percentages that seem to refer to distinct aspects (i.e.46%, 26% and 15.5%) provide an overall percentage of 82% -I cannot see how such a small percentage for each kind of annotations compose such a vast overall percentage.I perceive the same issue also for the comparison between RelTopic and the human annotators.Which particular approach has it been used to compute such percentages?Did the authors use some well-known statistics?If only simple percent agreement calculation has been used, why has this been preferred against more robust statistics such as Cohen's kappa and Fleiss' kappa, proposed for these purposes?Full availability of data