User-Defined Aggregates for Datamining.
Haixun Wang, Carlo Zaniolo · 1999
User-defined aggregates can be the linchpin of sophisticated datamining functions and other advanced database applications. This is demonstrated by our efficient implementation on DB2 of SQL3 user-defined aggregates extended with early returns, which we have used to implement several data mining algorithms. Aggregates with early returns are monotonic and can thus be used freely in recursive queries. 1 Introduction The importance of rollups and data cubes in decision support is universally recognized, and these aggregates have now become sine qua non built-ins in the new releases of commercial DBMSs. Yet, we claim that both researchers and vendors have overlooked UserDefined Aggregates (UDA)s, which can play an even more critical and pervasive role in database-centric data mining applications. Indeed, we show that: ffl Many data mining algorithms rely on specialized aggregates. ffl The number and diversity of these aggregates implies that (rather than vendors adding ad hoc builtins, ...