Cache-efficient implementation and batching of tridiagonalization on manycore CPUs
Shuhei Kudo, Toshiyuki Imamura · 2019
We herein propose an efficient implementation of tridiagonalization (TRD) for small matrices on manycore CPUs. Tridiagonalization is a matrix decomposition that is used as a preprocessor for eigenvalue computations. Further, TRD for such small matrices appears even in the HPC environment as a subproblem of large computations.