Task aware machine learning for objective-informed algorithm design

Yinsong Wang · 2024

The task-agnostic design of machine learning algorithms can lead to practical drawbacks, such as data-hungry learning schemes, poor generalization performance, and expensive computation overhead. In this thesis, we address these challenges by developing algorithms inspired by the learning objective. Our focus encompasses three general problem classes: kernel learning, dynamic density estimation, and cross-modal learning. In the first part, we study the alignment of kernel trajectory with the learning objective using random features approximations. Several recent studies have explored data-dependent sampling of features, modifying the stochastic oracle from which random features are sampled. While proposed techniques in this realm improve the approximation, their suitability is often verified on a single learning task. Here, we propose a task-specific scoring rule to select random features, which can be employed for different applications with some adjustments. We restrict our attention to Canonical Correlation Analysis (CCA), and we provide a novel, principled guide for finding the score function maximizing the canonical correlations. We prove that this method, called ORCCA, can outperform (in expectation) the corresponding kernel CCA with a default kernel. Next, we consider random features approximation in a multi-agent regression scheme. We present a fully distributed estimation algorithm where agents exchange local estimates with their neighbors to collectively identify the true value of the locally partially-observable parameter. This distributed update provably provides an asymptotically unbiased estimator of the unknown parameter, i.e., the first moment of the expected global error converges to zero asymptotically. We further analyze the efficiency of the proposed estimation scheme by establishing an asymptotic upper bound on the variance of the global error. In the second part, we designsliding-window" kernel density estimators for dynamic density tracking using objective information. Real-time density estimation is ubiquitous in many applications, including computer vision and signal processing. Kernel density estimation is arguably one of the most commonly used density estimation techniques, and the use of "sliding window" mechanism adapts kernel density estimators to dynamic processes. First, we derive the asymptotic mean integrated squared error (AMISE) upper bound for thesliding window'' kernel density estimator. This upper bound provides a principled guide for devising a novel estimator, which we name the temporal adaptive kernel density estimator (TAKDE). Compared to heuristic approaches for "sliding window" kernel density estimator, TAKDE is theoretically optimal in terms of the worst-case AMISE. Second, we consider the exact mean integrated squared error (MISE) for evolving Gaussian density estimation. We provide a principled guide for choosing the optimal weight sequence by theoretically characterizing the exact MISE, which can be formulated as a constrained quadratic programming. Our proposed weight sequence is shown to be theoretically optimal and empirically better than existing heuristic approaches. In the third part, we study the role of the learning objective in cross-modal learning with missing information (missing modality and missing label). First, we explore improving generalization performance in cross-modal imputation by introducing measure consistency regularization into the learning objective. We propose a unified framework, namely neural distance penalty, which enforces a measure consistency between observed data and imputed data by incorporating a neural net distance into the learning objective. We theoretically show the model-agnostic, distribution-agnostic advantage of neural distance penalty in improving the estimation error of the imputation function using the partially observed data. Second, we investigate a cross-modal learning framework, where the objective is to enhance the performance of supervised learning in the primary modality using an unlabeled, unpaired secondary modality. Taking a probabilistic approach for missing information estimation, we show that the extra information contained in the secondary modality can be estimated via Nadaraya-Watson (NW) kernel regression, which can further be expressed as a kernelized cross-attention module (under linear transformation). This expression lays the foundation for introducing The Attention Patch (TAP), a simple neural network add-on that can be trained to allow data-level knowledge transfer from the unlabeled modality.--Author's abstract

Read the paper · More papers on PaperTik