Efficiently support MapReduce-like computation models inside parallel DBMS

Qiming Chen, Andy Therber, Meichun Hsu, Hans Zeller, Bin Zhang, Wu Ren · 2009

While parallel DBMSs do support large scale parallel query processing on partitioned data, the reach of more general applications relies on User Defined Functions (UDFs). However, the existent UDF technology is insufficient both conceptually and practically. A UDF is not a relation-in, relation-out operator, which restricts its ability to model complex applications defined on a set of tuples rather than on a single one, and to be composed with other relational operators in a query. Further, to interact with the query execution efficiently, a UDF must be coded with complex interactions with DBMS internal data structures and system calls which is often beyond the expertise of an analytics application developer.

Read the paper · More papers on PaperTik