Data mining and compression : where to apply it and what are the effects?

Phillip Taylor, Nathan Griffiths, Xu Zhou, Alexandros Mouzakitis · Warwick Research Archive Portal (University of Warwick) · 2019

In data mining it is important for any transforms made to training data to be replicated on evaluation or deployment data. If they is not, the model may perform poorly or be unable to accept the input. Lossy data compression has other considerations, however, for example it may not be known whether or not lossy compression will be applied to deployment data, or if a variable compression ratio is to be used. Furthermore, lossy data compression typically reduces noise, which may not affect or even improve model performances, and performing feature selection on lossy data may find better features than selecting from the original data. In this paper, we investigate the effects of selecting features, learning, and making predictions from data that has been compressed using lossy transforms. Using vehicle telemetry data, we determine where in the data mining methodology lossy compression is detrimental or beneficial, and how it should be compressed. We also propose a specialised feature selection approach that considers predictive performance alongside compressibility, measured by compressing them either individually or in a single concatenated stream

Read the paper · More papers on PaperTik