Load Shedding Techniques for Data Stream Systems

Brian Babcock, Mayur Datar, Rajeev Motwani · 2003

Many data stream sources (communication network traffic, HTTP requests, etc.) are prone to dramatic spikes in volume. Because peak load during a spike can be orders of magnitude higher than typical loads, fully provisioning a data stream monitoring system to handle the peak load is generally impractical. Therefore, it is important for systems processing continuous monitoring queries over data streams to be able to adapt to unanticipated spikes in input data rates that exceed the capacity of the system. An overloaded system will be unable to process all of its input data and keep up with the rate of data arrival, so load shedding, i.e., discarding some fraction of the unprocessed data, becomes necessary in order for the system to continue to provide up-to-date query responses. While some heuristics for load shedding have been proposed earlier ([C02, M03]), a systematic approach to load shedding with the objective of maximizing query accuracy has been lacking. The main contributions of our work are:

Read the paper · More papers on PaperTik