Efficient data inventory management

Athanassios Papadimitriou, John W. Mamer · Medical Entomology and Zoology · 2004

This dissertation attempts to provide insight to the problem of efficient data management in a decision making context. It consists of two parts. The first part discusses the datagraph model and its offspring, the SESAME language and system. The datagraph extends relational algebra and sets the foundation upon which SESAME can be used as middleware to broaden the capabilities of an off-the-shelf Relational Database Management System (RDBMS). While classic relational algebra is basic in semantics only relating entities via key/foreign key relationships, the datagraph is rich in mathematical semantics via the use of hyperedges that succinctly establish functional mappings between entities. We use SESAME to model a simplified what-if analysis portfolio management application for a brokerage house and test its performance against a standard RDBMS. The results are encouraging. From a semantic modeling point of view, we are able to model complex situations, far beyond the capabilities of classic relational algebra and traditionally served by the flexible but limited in scalability and reusability spreadsheets. From a performance point of view, we find that for a set of common what-if queries, SESAME's substitution and rewriter modules improve total execution time performance by an order of magnitude or more. The second part of this dissertation aims at the same fundamental goal as the first part, namely, the efficient management and use of information. Yet, in the second part, this goal is pursued in a different fashion. While the theoretical foundation of the first part was database theory and—in particular—relational algebra, the second part makes use of stochastic modeling and a choice of optimization techniques commonly found in the operations literature. Data, stripped from the technicalities of the database world, is treated very much like a physical good in an operations-like fashion. Data is perishable, substitutable and is used to meet demand at a holding and maintenance cost. More importantly, it needs to be periodically replenished. We develop a series of stochastic optimal update policies that prescribe when to “order” new data (update the database). These policies weigh the cumulative decision error due to out-of-date data against the cost of performing a data update assuming an infinite horizon. Not surprisingly, some of the results we obtain in this novel context are not very different from established counterparts in inventory theory.

Read the paper · More papers on PaperTik