High performance implementation of tridiagonalization on the SR8000

Ken Naono, Y. Yamamoto, Mitsuyoshi Igai, Hiroyuki Hirayama · 2000

The methods of high performance tridiagonalization on the HITACHI SR8000 are described and evaluated. To achieve high performance, we adopted the blocked tridiagonalization and the scattered square decomposition. In addition, to achieve more performance in one node, we took the ways of the rectangular computation in the diagonal blocks and the loop integration for reducing the number of read/write operations. On one node of the SR8000, we achieved about 4.0 Gflop/s in the 4000-dimension tridiagonalization of a real symmetric matrix. This is much better than the 2.9 Gflop/s of our matrix library's on the HITACHI S-3800, which has the same peak performance with one node of the SR8000.

Read the paper · More papers on PaperTik