On Compression for the Complete IoT Data Lifecycle:From Collection to Analytics
Marcell Fehér · 2023
As the global volume of data is growing exponentially, our digital infrastructure is under tremendous stress. Experts forecast an over 4x increase in the amount of generated, transmitted, stored and processed data by 2025 compared to the 2018 levels, primarily fueled by Internet-connected sensors and devices commonly referred to as the Internet of Things (IoT). This trend has accelerated several research and development fields that aim to cope with this unprecedented growth, including the subject of our focus, data compression. It is versatile and allows for improvements at all stages of the IoT data lifecycle. Compressed data takes up less memory and storage on end devices and requires less energy to transmit. Efficient compression schemes significantly reduce the long-term storage requirements of data warehouses and cloud storage. It can even speed up analysis by lowering the performance penalty of fetching data from hard drives, utilizing main memory more efficiently and even allowing direct query evaluation. This thesis explores how Generalized Deduplication (GD), a novel data compression framework enables new possibilities from data transmission and privacy to efficient querying. To investigate source compression, we studied the use case of smart metering, where connected devices measure electricity consumption and report it to the utility provider at configurable time intervals. We have proposed four new compression methods that considerably outperform the currently used algorithms, allowing an order of magnitude more frequent reporting simultaneously. Evaluations were conducted on a real-life dataset including 95 Danish consumers over nine months. We have found that one of our techniques, which exploits redundancies in the data format, reduces message size by up to 83% when readings are immediately uploaded. Our pattern-based compressor, which is applicable for a wide range of multidimensional time series well beyond smart meter readings, achieves compression gains up to 74% when power consumption is reported every hour. Data compression may have unintended effects when applied to sensitive information, as demonstrated by multiple side-channel attacks on compressed-encrypted data transfer. Since power consumption can be an indicator of the habits of the people living in the same household, privacy concerns must be considered when uploading readings. To this end, we first quantified the correlation between compressed message size and encoded power use and found that some of the standard compressors exceed 75%, leaking enough information that allows a passive observer of the encrypted network traffic to infer the daily routine of tenants. Based on this exploit, we presented a possible large-scale attack that allows easy classification of compromised households by similar daily routines, using acquired network logs. We have shown that the message sizes of our novel data compressors are less than 20% correlated with the power consumption when reporting period is under an hour and 35% more resistant to the presented clustering-based attacks than the currently used methods. IoT data is often analyzed in relational databases using complex queries. Many systems support multiple compression schemes to reduce storage requirements and speed up loading data from high-latency hard drives. We have designed new integer compressors based on Generalized Deduplication and used Hyrise, an open-source columnar database for evaluations with industry-standard analytical benchmarks, TPC-H and Join Order Benchmark. Our results show that GD-based encoders can achieve 19% higher compression than LZW, complete the benchmark 24x faster, and compress 48% better than the default Dictionary Encoding while being only 20% slower in queries. Our novel segments represent a new breed of encoders, as they can adapt their configuration to both the input data and the experienced workload as data becomes less relevant over time. A self-driving database that detects a significant change of access and query patterns can instruct our compressors to adjust data representation and optimize it for the updated priorities. With an order-preserving transformation function, GD-compressed data segments can even be searched without reconstructing a single value, thanks to our proposed novel algorithm.