Diag-Join: An Opportunistic Join Algorithm for 1:N Relationships

Sven Helmer, Till Westmann, Guido Moerkotte · 1998

Abstract Time of creation is one of the predominant (often implicit) clustering strategies found not only in Data Warehouse systems: line items are created together with their corresponding order, objects are created together with their subparts and so on. The newly created data is then appended to the existing data. We present a new join algorithm, called DiagJoin, which exploits time-of-creation clustering. If we are able to take advantage of timeof-creation clustering, then the performance evaluation reveals the superiority of Diag-Join over standard join algorithms like block-wise nested-loop join, GRACE hash join, and index nested-loop join. We also present an analytical cost model for Diag-Join. 1 Introduction During the evaluation of queries in Data Warehouses, relations containing millions or even billions of tuples need to be joined. Joins involving fact tables are very costly operations. Evidently, fast join algorithms are very important in this environment.

Read the paper · More papers on PaperTik