Application des grandes matrices aléatoires aux séries temporelles multivariables

Daria Tieplova · HAL (Le Centre pour la Communication Scientifique Directe) · 2020

A number of recent works proposed to use large random matrix theory in the context of high-dimensional statistical signal processing, traditionally modeled by a double asymptotic regime in which the dimension of the time series and the sample size both grow towards infinity. These contributions essentially addressed detection or estimation schemes depending on functionals of the sample covariance matrix of the observation. However, fundamental high-dimensional time series problems depend on matrices that are more complicated than the sample covariance matrix. The purpose of the present PhD is to study the behaviour of the singular values of 2 kinds of structured large random matrices, and to use the corresponding results to address an important statistical problem. More specifically, the observation (y_n)_{nin Z} is supposed to be a noisy version of a M-dimensional time series (u_n)_{nin Z} with rational spectrum that has some particular low rank structure, the additive noise (v_n)_{nin Z} being an independent identically distributed sequence of complex Gaussian vectors with unknown covariance matrix. An important statistical problem is the estimation of the minimal dimension P of the state space representations of u from N samples y_1,.., y_N. If L is any integer larger than P, the traditional approaches are based on the observation that P coincides with the rank of the autocovariance matrix R^L_{f|p} between the ML-dimensional random vectors (y_{n+L}^T,..,y_{n+2L-1}^T)^T and (y_{n}^T,.., y_{n+L-1}^T)^T, as well as with the number of non zero singular values of the normalized matrix C^L = (R^L)^{-1/2}R^L_{f|p} (R^L)^{-1/2} where R^L represents the covariance matrix of the above ML-dimensional vectors. In the low-dimensional regime where N->+infty while M and L are fixed, the matrices R^L_{f|p} and C^L can be consistently estimated by their empirical counterparts hat{R}^L_{f|p} and hat{C}^L, and P can be evaluated from the largest singular values of hat{R}^L_{f|p} and hat{C}^L. If however M and N->+infty in such a way that ML/N converges towards 01/2} and that the structure of S_R is more intricate. It is moreover established that all the singular values of hat{R}^L_{f|p} and hat{C}^L are located in the neighbourhood of S_R and S_C respectively. When u is present, the low rank structure of u is used in order to study whether some singular values of hat{R}^L_{f|p} and hat{C}^L escape from S_R and S_C} It is shown that the number of singular values of hat{R}^L_{f|p} located outside S_R is not directly related to P, while, fortunately, P coincides with the number of singular values of hat{C}^L that are larger than 2sqrt{c*(1-c*)}, provided c*<1/2, the signal u is powerfull enough compared to the noise and the non zero singular values of C^L are large enough. These results imply that while the singular values of hat{R}^L_{f|p} can be used in order to estimate P consistently in the standard low-dimensional regime, this is no longer the case in the high-dimensional context considered here. Fortunately, under certain assumptions, P canstill be consistently estimated as the number of singular values of hat{C}^L that are larger than 2 sqrt{c*(1-c*)}

Read the paper · More papers on PaperTik