Enabling data-efficient trajectory prediction via coreset selection
Ruining Yang · 2024
The safety and reliability of autonomous vehicles depend on accurate trajectory predic- tion. Although deep learning models have achieved excellent performance in trajectory prediction, most of them rely on large-scale training data. The training process requires extremely high compu- tational efforts. However, not all data samples are equally important for model training. To address the issue of high training costs, we propose a new data-efficient training method based on coreset selection, suitable for trajectory prediction tasks. Our method focuses on selecting a small subset of trajectory data that significantly contributes to the model. Despite significantly reducing the amount of data, we are able to maintain the accuracy of the model, and the selected subset can be used to train other trajectory prediction models. Our method is divided into multiple steps. First, the data is pre-processed. Based on the number of agents in the current scenario, we filter out data samples with fewer agents in the scene. We then calculate the submodular gain for each data sample dur- ing the dataset selection phase to identify and retain the data that are most beneficial to the model. Finally, we use the optimized subset to train the trajectory prediction model. Experimental results show that the subset of trajectory data selected based on the coreset selection method can not only accelerate training and maintain the accuracy of the model but also demonstrate excellent general- ization capabilities, especially when the amount of data is small. Our method is the first to apply a data-efficient training approach to the trajectory prediction task, providing an insightful solution to improve the training efficiency of autonomous driving systems.--Author's abstract