Quantile estimation in dynamic and stationary environments using the theory of stochastic learning
Anis Yazidi, Hugo L. Hammer · ACM SIGAPP Applied Computing Review · 2016
The goal of our research is to estimate the quantiles of a distribution from a large set of samples that arrive sequentially. Since the data set is large, the model we choose is that the data cannot be stored, but rather that estimates of the quantiles are computed in a real-time setting. In such settings, classical estimators that require storing the whole history of the data (or stream) cannot be deployed. In this paper, we present an incremental quantile estimator of a distribution, i.e., one that utilizes the previously-computed estimates and only resorts to the last sample for updating these estimates. The state-of-the-art work on obtaining incremental quantile estimators is due to Tierney [12], and is based on the theory of stochastic approximation. However, a primary shortcoming of the latter work is the requirement to incrementally build local approximations of the distribution function in the neighborhood of the quantiles. This requirement, unfortunately, increases the complexity of the algorithm. In addition to treating the case of a constant update parameter, we extend our work to include the case of a decreasing update parameter. Such modification is suitable for the case of a stationary environment where the true quantile is invariant over time. Experimental results demonstrate that our estimator outperforms the state-of-the-art estimators. In addition, it also copes with dynamic environments.