Clustering Multi-Blocs et Visualisation Analytique de Données Séquentielles Massives Issues de Simulations du Véhicule Autonome

Étienne Goffinet · HAL (Le Centre pour la Communication Scientifique Directe) · 2021

Advanced driver-assistance systems (ADAS) development remains one of the biggest challenges car manufacturers must tackle to provide safe driverless cars. The reliable validation of these systems requires assessing their reaction's quality and consistency in a broad spectrum of driving scenarios. In this context, Groupe Renault uses large-scale simulation systems, which accurately reproduce the physical driving conditions and produce large quantities of high-dimensional time series data. The role of the ADAS developer is to explore these datasets to determine precisely the capabilities of the driving assistance system under test and, if necessary, to refine its design.A simulated dataset can contain up to several hundreds of thousands of simulations for a given use case, described by several hundreds variables. With datasets of such dimensions, the expert's work requires a great deal of field knowledge and involves a time-consuming exploration process.This thesis objective is to produce algorithms and tools to help explore, structure and analyze these massive simulation datasets. We propose four probabilistic approaches for time series clustering, each based on a specific datasets structure hypothesis.The first contribution of this thesis is built on a scenario-based construction hypothesis and performs a dictionary-based classification of univariate time series. This method uses an existing piecewise polynomial model to segment the time series independently, then constructs a driving pattern dictionary. This dictionary is used to recode the time series into chains of categories, which are then partitioned with a hierarchical clustering algorithm. The subsequent contributions focus on the analysis of multivariate datasets with multi-blocks models that simultaneously discriminate the simulations and the variables based on their distributions and row-partitions. These Bayesian Non-Parametric models natively integrate a model selection, which enables automatic model resizing with respect to the datasets contents.Finally, the usefulness of these contributions is illustrated with several applications on industrial use cases: the validation of emergency braking, lane-keeping and emergency avoidance systems.

Read the paper · More papers on PaperTik