An Overview of Data
Adam P. Tashman · 2024
Data is information that can be stored and used by a computer. We study various aspects of data in this chapter, including data type, volume, structure, and velocity. These properties help define how data can be stored, processed, and used. We study two important approaches for data ingestion: API and web scraping. Data integration consists of preparing and combining the ingested data. The common pipeline consists of extracting, transforming, and loading the data (the ETL pipeline). Metadata is data about data, such as movie genre. Data can be at rest, in use, or in motion. The velocity of the data is the speed at which it enters a system for processing. Some datasets are finite and they can be processed in a large chunk. Data that is infinite is called streaming data. Data at rest may be structured, semi-structured, or unstructured. Structured data can be stored in tables. Semi-structured data is generally hierarchical, and it is stored as key-value pairs. Unstructured data comes from a variety of sources, such as audio files. To be useful, data needs to be credible and reliable, and it must faithfully represent the population. When data is not representative, there will be bias.