Handling Data Skew in Multiprocessor Database Computers Using Partition Tuning

Kien A. Hua, Chiang Lee · Very Large Data Bases · 1991

Shared nothing multiprocessor archit.ecture is known t.o be more scalable to support very large databases. Compared to other join strategies, a hash-ba9ed join algorithm is particularly efficient and easily parallelized for this computation model. However, this hardware structure is very sensitive to the data skew problem. Unless the parallel hash join algorithm includes some load balancing mechanism, skew effect can deteriorate t.he system performance

Read the paper · More papers on PaperTik