Exactly Decoding a Vector through Relu Activation
Samet Oymak, Muhammad Salman Asif · 2019
We consider learning a d-dimensional parameter w through nonlinear input/output relation governed by ReLU activation. We study a supervised learning setup in which we want to decode w from input/output pairs (x, y). We consider an additive model with nonlinear ReLU activation that can be represented as y = Σk=1dReLU(w[k] + x[k]). Such a model appears in representation learning and recommendation systems where w corresponds to an unknown embedding of a user or item and the x correspond to embedding of known probe vectors. In this paper, we show that a gradient descent algorithm linearly converges with O(d) samples and quickly finds the true parameter w under mild assumptions. Our assumptions are in terms of the input distribution that captures the fundamentals of the problems. We also demonstrate the performance of our algorithm with numerical simulations.