Prediction of driver actions with long short-term memory recurrent neural networks
Martin Torstensson · Chalmers Publication Library (Chalmers University of Technology) · 2018
The most prominent factor behind traffic accidents today is human errors.There are many ways, in which problematic behaviors such as inattention can be mitigated.One of the tools used for this purpose is warning systems.These warning systems needs to give warnings both at a time where the driver needs it as well as give the driver a sufficient window to react.In order to achieve these goals information of the state of the driver can be vital to know how curtain situations should be handled.A prediction of future events could be used in order to increase the amount of time between the warning and the dangerous event.This report explores possibilities of using recurrent neural networks with long short-term memory for prediction of eight different driver actions inside of a vehicle, such as glancing and reaching inside of the vehicle among others.Other studies in the domain of predictions of driver actions have had the focus on car movements such as breaking and lane changing, while this study concentrates on the state of the driver.There is potential for these predictions to improve a warning system and give a driver more time to react to a given situation.The predictions are based on sequences of actions, which are generated from sequences of images with a convolutional neural network.A dataset consisting of sequences of images used in the report was gathered at RISE Viktoria AB.The hyperparameters of the recurrent neural network, such as the number of hidden units and amount of layers, was chosen with Bayesian optimization.An addition of a parallel input of optical flow created from the input images was found to improve the performance of the convolutional neural network.The complete network achieved an average prediction accuracy of 80% for the next frame predictions and 62% after 20 frames.A comparison where the predictions were set to the last element in the input achieved an accuracy of 79% for one frame ahead and 49% after 20 frames.