A Data Miner Looks at SQL

Gordon S. Linoff · 2015

The focus on data mining has historically been on complex algorithms developed by statisticians and machine-learning specialists. This chapter presents a synopsis of the key concepts discussed in the various chapters of this book. It introduces SQL and relational databases such as Hadoop, Hive, NoSQL, from different perspectives important for data mining and data analysis. The first perspective is the structure of the data, with a particular emphasis on entity-relationship diagrams. The second is the processing of data using dataflows. The third, and strongest thread through subsequent chapters, is the syntax of SQL itself. The focus in this chapter, and throughout the book, is on SQL for querying. The important functionality of SQL and how it is expressed is discussed, with particular emphasis on joins, group bys, and subqueries, because these play an important role in data analysis.

Read the paper · More papers on PaperTik