Should co‐occurrence data be normalized? A rejoinder
Loet Leydesdorff · Journal of the American Society for Information Science and Technology · 2007
The ArgumentOur argument was analytical, and not only based on the visualizations.For example, in the dataset under study, Ahlgren, Jarneving, and Rousseau (2003, p. 556) found a correlation of r ϭ ϩ0.74 between "Schubert" and "Van Raan," while Leydesdorff and Vaughan (p.1620) report r ϭ Ϫ0.131 (p Ͻ 0.05) using the underlying (and asymmetrical) citation matrix.The difference can be explained by the generation of spurious correlations between the co-cited items.One can calculate from Table 7 of Ahlgren et al. (2003, p. 555) that the margintotal of co-citations of these two authors was 139 and 138, respectively.However, these numbers are based on only 60 and 50 citations, respectively.In other words, these authors are heavily co-cited with their collaborators in this network, and thus the total number of co-citations for them is more than twice the number of citations.This "double counting" of co-citations because of underlying coauthorship relations generates spurious correlations in the co-occurrence matrix.The attribute data in the asymmetrical citation matrix contain more information than the symmetrical co-occurrence matrix.The latter can be generated from the former, but an attribute matrix cannot be generated from a co-occurrence matrix because information is lost during the transformation.1 Waltman and Van Eck (2007; hereafter W&vE) did not distinguish sufficiently between these two matrices when they formulated: "In order to correct for the differences in the number of times authors are cited, co-citation matrices should be normalized, for example using the Pearson correlation."It does not follow from the differences in the number of times authors are cited, that cocitation matrices should be normalized.If one wishes to show the similarities between authors, author attributes should be normalized (Burt, 1982; Schneider & Borlund, in press).In the Web environment, the approach of retrieving first citation data is often not feasible.In that case, the normalization of co-citation matrices should not be based on the Pearson correlation or the cosine but on the Jaccard index (Luukkonen, Tijssen, Persson, & Sivertsen, 1993;Small, 1973) or on probabilistic measures, which normalize observed cell values against expected ones (Michelet, 1988;Zitt, Bassecoulard, & Okubo, 2000).Unlike the Pearson correlation or the cosine, the latter measures do not use the information contained in the distributions but only the cell values and the margintotals of the subsets (Leydesdorff, 2007).