Learning with subsampled kernel-based methods: Environmental and financial applications
M. Aminian Shahrokhabadi, A. Neisy, Emma Perracchione, Mirko Polato · Padua Research Archive (University of Padova) · 2019
Kernel machines are widely used tools for extracting features from given data.In this context, there are many available techniques that are able to predict, within a certain tolerance, the evolution of time series, i.e. the dynamics of the considered quantities.However, the main drawback is that measurements are usually affected by noise/errors and might have gaps.For instance, gaps might be due to several problems of the physical instruments that produce measurements.In these cases the learning and prediction steps for capturing the trend of time series become very hard.To alleviate these difficulties, we construct a primary kernel-based approximant, which is indeed a model, with the double aim: to fill the gaps and to filter noisy data.The so-constructed smoothed samples are used as training sets for a kernel-based online model.We claim that the subsampled training phase makes the predicted results more stable.Applications to real data for environmental and financial observations support the validity of our results.https://github.com/makgyver/vlabtestrepo/.The guidelines of the paper are as follows.In Section 2, we briefly review the basic aspects of RBF theory, focusing on the so-called reproducing kernel property.Section 3 presents our scheme for RBF regression.Such algorithm is tested via extensive numerical experiments in Section 4. The last section deals with conclusions and future work. Reproducing kernel propertyIn this section we review the main theoretical aspects of kernel-based methods.For further details refer to the books [4,9,10,23].Let us suppose that we have X N = {x i , i = 1, . . ., N } ⊂ Ω, a set of distinct data points (or data sites or nodes), arbitrarily distributed on a domain Ω ⊆ d , and an associated set F N = { f i = f (x i ), i = 1, . . ., N } of data values (or measurements or