Config 2.0

Jawad Tahir · 2021

Distributed stream processing systems (DSPSs) have enabled us to build scalable, fast, and reactive applications. As faults are common in any distributed system, DSPSs allow for recovering from faults using various techniques, for example, snapshotting their state on a key-value (KV) store and then replaying the events in case of a failure. The performance of the recovery process is dependent on the workload, configuration of the DSPS and performance of the KV store. Performance tuning of a KV store is a notoriously difficult job. In this thesis proposal, we propose a reinforcement learning (RL) based agent which configures the DSPS and the KV store under varying workloads to minimize the impact of a failure on the performance.

Read the paper · More papers on PaperTik