Relating information complexity and training in deep neural networks

Alex Gain, Hava Siegelmann · 2019

Deep Neural Networks may be costly to train, and if testing error is too large, retraining may be required, unless lifelong learning methods are applied. Crucial to addressing learning at the edge, without access to powerful cloud computing, is the notion of problem difficulty for non-standard data domains. While it is known that training is harder for classes that are more entangled, the complexity of data points was not previously studied as an important contributor to training dynamics. We analyze data points by their information complexity and relate the complexity of the data to the test error. We elucidate training dynamics of DNNs, demonstrating that high complexity datapoints contribute to the error of the network, and that training DNNs consist of two important aspects - (1) Minimization of error due to high complexity datapoints, and (2) Margin decrease where entanglement of classes occurs. Whereas data complexity may be ignored when training in a cloud, it must be considered as part of the setting when training at the edge.

Read the paper · More papers on PaperTik