Distance Learning for Analog Methods (1st revision)
Paul Platzer, Arthur Avenas, Bertrand Chapron, Lucas Drumetz, Alexis Aurelien Mouche, Pierre Tandeo, Léo Vinour · Institutional Archive of Ifremer (French Research Institute for Exploitation of the Sea) · 2025
Analogs are similar states of a system, occurring at remote times within independent numerical simulations or previous observations. This concept has emerged in atmospheric sciences and was further used in ocean sciences for forecasting, downscaling, upscaling, and extreme event attribution. The distance used to find and rate analogs is a key feature of analog methods. Most studies are based on the Euclidean distance or other pre-defined metrics. In this investigation, we adapt distance learning algorithms originally designed for classification and regression to statistical forecasting objectives, using the continuous-ranked probability score as a loss function. Our algorithm allows to jointly optimize three key hyperparameters of analog methods: the feature space, distance, the number of analogs used. In particular, this algorithm allows to reduce the feature space dimension while preserving high performances, a key requirement for small-size datasets. We test our algorithm on an idealized chaotic system and a tropical cyclone dataset. These experiments suggest that the optimal distance depends on the forecast horizon and the number of available data, and that our algorithm allows for reasonable performances of analog ensemble methods even for small-size datasets. Our algorithm runs faster then existing grid-search-like analog distances optimization algorithms. This allows to test and optimize a wider class of distances, for instance weighting a large number of predictive variables. Our approach is not limited to forecasting and can assist the search for optimal hyperparameters of any analog method, enhancing exploration possibilities and improving overall performances.