Another step toward demystifying deep neural networks

Michael Elad, Dror Simon, Aviad Aberdam · Proceedings of the National Academy of Sciences · 2020

The field of deep learning has positioned itself in the past decade as a prominent and extremely fruitful engineering discipline. This comeback of neural networks in the early 2000s swept the machine learning community, and soon after found itself immersed in practically every scientific, social, and technological front. A growing series of contributions established this field as leading to state-of-the-art results in nearly every task, recognizing image content, understanding written documents, exposing obscure connections in massive datasets, facilitating efficient search in large repositories, translating languages, enabling a revolution in transportation, revealing new scientific laws in physics and chemistry, and so much more. Deep neural networks not only solve known problems but offer, in addition, unprecedented results in deploying learning to problems that until recently were considered as hopeless or only weakly successful. These include automatically synthesizing text–media, creating musical art pieces, synthesizing realistic images and video, enabling competitive game-playing, and this list goes on and on. Amazingly, all these great empirical achievements are obtained with hardly any theoretical foundations that could provide a clear justification for the architectures used, an understanding of the algorithms that accompany them, a clear mathematical reasoning behind the various tricks employed, and above all, the impressive results obtained. The quest for a theory that could explain these ingredients has become the Holy Grail of data sciences. Various impressive attempts to provide such a theory have started to appear, relying on ideas from various disciplines (see, e.g., refs. 1⇓⇓⇓⇓⇓⇓⇓–9). The paper by Papyan et al. (10) in PNAS adds an important layer to this vast attempt of developing a comprehensive theory that explains the behavior of deep learning solutions. In this Commentary, we provide a wide context to their results, highlight and clarify their contribution, and raise … [↵][1]1To whom correspondence may be addressed. Email: elad{at}cs.technion.ac.il. [1]: #xref-corresp-1-1

Read the paper · More papers on PaperTik