An Autonomous Decentralized Architecture for Distributed Data Management and Dissemination

M. Brian Blake, Patricia A. Liguori · IEICE Transactions on Information and Systems · 2001

Over recent years, “Internet-able” applications and architectures have been used to support domains where there are multiple interconnected systems that are both decentralized and autonomous. In enterprise-level data management domains, both the schema of the data repository and the individual query needs of the users evolve over time. To handle this evolution, the resulting architecture must enforce the autonomy in systems that support the client needs and constraints, in addition to maintaining the autonomy in systems that support the actual data schema and extraction mechanisms. At the MITRE Corporation, this domain has been identified in the development of a composite data repository for the Center for Advanced Aviation System Development (CAASD). In the development of such a repository, the supporting architecture includes specialized mechanisms to disseminate the data to a diverse evolving set of researchers. This paper presents the motivation and design of such an architecture to support these autonomous data extraction environments. This run-time configurable architecture is implemented using web-based technologies such as the Extensible Markup Language (XML), Java Servlets, Extensible Stylesheets (XSL), and a relational database management system (RDBMS). 1.0 Introduction and Motivation At the MITRE Corporation-Center for Advanced Aviation System Development (CAASD), researchers develop simulations for both design-time and real-time analysis. This research constitutes a wealth of knowledge in the area of air traffic management and control. This division of MITRE is split into a large number of individual groups that investigate various problems comprising the air traffic domain. Although the groups analyze different problems, the data to support the investigations are typically the same. Also, these individual groups develop simulations that require the data in different formats (i.e. specialized text files with delimited data, database format, XML, etc.) Moreover, each group looks at different subsets of data that may cross multiple data sources. Researchers are currently provided with data from outside sources that is gathered and distributed by a data librarian. This data is usually distributed in the same media and format in which it is obtained. The CAASD Repository System (CRS) team at MITRE has identified the need for obtaining the desired raw data from outside sources and building a composite data repository that serves the need of this diverse environment. This paper presents the architecture that will allow this data extraction not only for the air traffic domain but also for other domains throughout MITRE such as army strategic motions and telemetry. The goal of the CRS team is to develop a dynamic architecture that will facilitate this data extraction for any data schema with the inclusion of some specific meta-data. In order to ensure the dynamic nature of this architecture, the CRS team takes a distributed web-based approach that separates the major components of the architecture into various autonomous modules. Though the goal of the team is toward an architecture that will not change, this autonomy will ensure the reusability of each component in the architecture. The next section this paper provides an overview of the CRS architecture and its autonomous modules. Section 3 provides a description of the actual technologies used to implement the architecture. In Sections 4 and 5, we discuss its autonomy and current usage. 2. The CRS Architecture The CRS architecture was devised to support a diverse set of customers/users. These customers internal to MITRE-CAASD use numerous technologies, programming languages, and interfaces. Web access is the one technology common to all groups. Therefore, the direction in designing the architecture was to use Internet technology as much as possible in gathering data request information and delivering the data to the customers. The CRS architecture is composed of four autonomous modules that fit seamlessly into the Internet paradigm. These modules are the Client Interface Module, the Interface Specification Module, the Presentation and Query Module, and the Database Extraction Module. These modules can be split across three layers, the Interface Layer, the Presentation Layer, and the Data Storage Layer. These three layers and the underlying modules are illustrated in Figure 1. Client Interface Module Interface Specification Module Presentation and Query Module Database Extraction Module Interface Layer Presentation Layer Data Storage Layer Figure 1. CRS Architecture by Layers. The Interface layer is the layer by which users can connect to the system. This layer consists of the Client Interface Module. Currently, Internet browsers implement the Client Interface Module. The customers use the browsers to connect to the system. In the future, this module might also include some stand-alone applications, which support data streaming. The Presentation Layer contains the Interface Specification Module and the Presentation and Query Module. Both of these modules include software services that provide a graphical user interface. The Interface Specification module allows the customers to customize their user interface to meet their specific needs. This is important considering the diverse data needs. The Presentation and Query module allows the customer to choose a standard or specialized interface in order to request data. Later, this module will need to be enhanced to explicitly allow the specification of business and domain logic. This module packages the information that will later be used in the Data Storage Layer. The Data Storage Layer contains functionality to maintain and extract data from some data repository. This layer consists of software services for extracting data from the relational database management system (RDBMS). Each module can further be decomposed into individual autonomous components. The decomposition of the modules is illustrated in Figure 3. As previously mentioned, the Client Interface Module currently contains Internet browsers that connect to the CRS system. The system provides two main functions for the users. In the first main function, a user can access the Interface Specification Module and design a personalized query form. This functionality is designed for users that need to execute repetitious personalized queries. The Interface Specification Component saves this personalized query form in a shared file system. The other function allows the user to access a personalized or standard query form within the Presentation and Query module and execute a query on the data repository. The User Interface component has access to the shared file system that stores the personalized and standard forms. The Data Extraction module is a service to both the Specification module and the Presentation and Query Module. The Data Access component accepts connections from the User Interface or Interface Specification component to satisfy internal or external services. 3. An Implementation of CRS System At MITRE-CAASD, the CRS team has implemented the CRS architecture using various Internet technologies. Figure 2 presents the implementation that supports the details in the original architecture diagram. The CRS implementation mainly uses Java-based technologies. The browsers in the Client Interface Modules connect to the Java Servlet-based components in both the Specification Module and the Presentation and Query Module. Both the User Interface component and Interface Specification component are implemented with Java Servlets. These Servlets are integrated with Java classes that fulfill the underlying query services. The Servlet in the Interface Specification module accepts information from the browsers in the Client Interface module in the form of an HttpServletRequest. This information can be parsed and used to generate the specifications for the HTML-based query form. This module stores this user preference information as an XML file in a shared file system location. In building this XML file, the module gathers database specific information using the Data Extraction Module. This information is coupled with the user preference information. Subsequently, this XML can be processed with a generic XSL file to dynamically generate the HTML-based query form. This XML file is the centerpiece of the architecture as it is the basis for the execution of the system. This file contains detailed information that allows the Presentation and Query Module to be generic. The benefit here is to allow outside sources to use the same XML format, and without the Interface Specification module, to have the ability to use the Presentation and Query module for data retrieval. The Servlet in the Presentation and Query module also receives an HttpServletRequest from the browsers. This module receives two independent messages. The first HttpServletRequest designates the particular standard or user-personalized query form to display. This personalized query form will be specific to each user.

Read the paper · More papers on PaperTik