A Data Miner Looks at SQL
Gordon S. Linoff · 2015
The focus on data mining has historically been on complex algorithms developed by statisticians and machine-learning specialists. This chapter presents a synopsis of the key concepts discussed in the various chapters of this book. It introduces SQL and relational databases such as Hadoop, Hive, NoSQL, from different perspectives important for data mining and data analysis. The first perspective is the structure of the data, with a particular emphasis on entity-relationship diagrams. The second is the processing of data using dataflows. The third, and strongest thread through subsequent chapters, is the syntax of SQL itself. The focus in this chapter, and throughout the book, is on SQL for querying. The important functionality of SQL and how it is expressed is discussed, with particular emphasis on joins, group bys, and subqueries, because these play an important role in data analysis.