Scalable feature learning

Quoc Viet Le · 2013

Over the past decade, machine learning has emerged as a powerful methodology that empowers autonomous decision making by learning and generalizing from examples. Thanks to machine learning, we now have software that classifies spam emails, recog- nizes faces from images, recommends movies and books. Despite this success, machine learning often requires a large amount of labeled data and significant manual feature engineering. For example, it is difficult to design algorithms that can recognize ob- jects from images as well as humans can. This difficulty is due to the fact that data are high-dimensional (a small 100x100 pixel image is often represented as 10,000 di- mensional vector) and highly-variable (due to many factors of transformations such as translation, rotation, illumination, scaling, viewpoint changes). To simplify this task, it is often necessary to construct features which are invariant to transformations. Features have become the lens that machine learning algorithms see the world. Despite its importance to machine learning and A.I., the process of construct- ing features is typically carried out by human experts and requires a great deal of knowledge and time, typically years. Even worse, these features may only work on a restricted set of problems and it can be difficult to generalize them for other do- mains. It is generally believed that automating the process of creating features is an important step to move A.I. and machine learning forward. Deep learning and unsupervised feature learning have shown great promises as methods to overcome manual feature engineering by learning features from data. However, these methods have been fundamentally limited by our computational abil- ities, and typically applied to small-sized problems. My recent work on deep learning and unsupervised feature learning has mainly focused on addressing their scalability, especially when applied to big data. In particular, my work tackles fundamental challenges when scaling up these algorithms by i) simplifying their optimization problems, ii) enabling model parallelism via sparse network connections, iii) enabling robust data parallelism by relaxing synchronization in optimization. The details of these techniques are described below Making deep learning simple: While certain classes of unsupervised feature learning algorithms, such as Independent Component Analysis (ICA), are effective in learning feature representations from data, they are difficult to optimize. To ad- dress this, we developed a simplified training method, known as RICA, by introducing a reconstruction penalty as a replacement for orthogonalization. Via a sequence of mathematical equivalences, we proved that RICA is equivalent to the original ICA op- timization problem under certain hyperparameter settings. The new approach how- ever has more freedom to learn overcomplete representations and converges faster. Our proof also shows connections between ICA and other deep learning approaches (sparse coding, RBMs, autoencoders etc.). Our algorithm, RICA, in…

Read the paper · More papers on PaperTik