Descriptive statistics and visualization of data from the R datasets package with implications for clusterability

Naomi C. Brownstein, Andreas Adolfsson, Margareta Ackerman · Data in Brief · 2019

The manuscript describes and visualizes datasets from the datasets package in the R statistical software, focusing on descriptive statistics and visualizations that provide insights into the clusterability of these datasets. These publicly available datasets are contained in the R software system, and can be downloaded at https://www.r-project.org/ , with documentation provided at https://stat.ethz.ch/R-manual/R-devel/library/datasets/html/00Index.html . Further information on clusterability is found in the companion to this article, To Cluster or Not to Cluster: An Analysis of Clusterability Methods ? ( https://doi.org/10.1016/j.patcog.2018.10.026 ). Brief descriptions and graphs of the variables contained in each dataset are provided in the form of means, extrema, quartiles, standard deviation and standard error. Two-dimensional plots for each pair of variables are provided. Original references to the data sets are included when available. Further, each dataset is reduced to a single dimension by each of two different methods: pairwise distances and principal component analysis. For the latter, only the first component is used. Histograms of the reduced data are included for every dataset using both methods.

Read the paper · More papers on PaperTik