Management of transient xml messages
Ashish Kumar Gupta, Dan Mircea Suciu · 2005
Today, a typical large enterprise has thousands of autonomous, independently deployed applications running within different divisions. These applications, however, need to communicate and coordinate with each other to perform business processes. These conflicting goals are achieved by messaging---exchange of transient messages. The volume of messages in such a scenario can easily exceed 100 MB/s. We assume these messages to be in XML because XML has been accepted as the standard for exchange of information over the Internet. In messaging, the messages are sent via an intermediary called the message broker. The primary functions of a message broker are to route and transform incoming messages. Essentially, the broker has a large set of rules (also called subscriptions) and for each incoming message it evaluates the conditions and transformations requested by each of the rules before it forwards the message. To achieve integration of the enterprise applications, we need to develop efficient techniques to implement the routing and transformation functionality of the message broker. There are three obstacles that make this a challenging problem: (a) large sized workloads at brokers, (b) complexity of the rules, and (c) high message throughput. We assume that the messages are in XML and the rules are in XQuery. In this dissertation, we take on all three of these challenges and describe techniques to implement routing and transformation functionality of the message broker efficiently. We achieve efficiency and scalability by sharing common computations. We have studied three subsets of XQuery, namely, (a) Linear Xpath Filters, (b) Xpath Filters with branches and predicates, and (c) XQuery (except function calls). For each subset, we identify the different types of computations that can be shared. We then devise techniques to share these computations and perform extensive experimental evaluation of these techniques on real, synthetic and benchmark datasets. We also propose two light weight techniques, Stream Views and Stream IndeX (SIX) to speed up parsing.