Chapter 1: What Is Imperfect Data?
Ronald K. Pearson · Society for Industrial and Applied Mathematics eBooks · 2020
The title of this book immediately raises the question of what imperfect data is, exactly? To provide a practical answer, Sec. 1.1 addresses the “mirror image” question, “what is perfect data?,” followed by brief descriptions in Sec. 1.4 of 10 ways real datasets can and regularly do depart from this working definition of perfection. Before launching into these discussions, Sec. 1.3 summarizes the main data types considered in this book, which is important background information since some of the data anomalies considered here are either type specific (e.g., outliers in numerical data or thin levels in categorical data) or manifested very differently in different data types (e.g., missing data in numerical versus text data). Also, Sec. 1.3 briefly describes some important data types that are beyond the scope of this book (e.g., image data and graphs). To provide concrete illustrations of the data types and anomalies that are considered here, five real datasets are described in Sec. 1.2. Further, because software is a practical necessity in acquiring and analyzing datasets of realistic size, Sec. 1.2 also introduces the two primary software environments used in this book (R and Python), with some of the reasons for these particular choices.