Large-scale Predictive Analytics in Vertica

Subrat Kumar Prasad, Arash Mohammadali Zadeh Fard, Vishrut Gupta, Jorge Martínez, Jeff LeFevre, Vincent Eric Xu, Meichun Hsu, Indrajit Roy · 2015

A typical predictive analytics workflow will pre-process data in a database, transfer the resulting data to an external statistical tool such as R, create machine learning models in R, and then apply the model on newly arriving data. Today, this workflow is slow and cumbersome. Extracting data from databases, using ODBC connectors, can take hours on multi-gigabyte datasets. Building models on single-threaded R does not scale. Finally, it is nearly impossible to use R or other common tools, to apply models on terabytes of newly arriving data.

Read the paper · More papers on PaperTik