Four‐perspective model for metadata requirements engineering
Jung Sun Oh, Sheila Denn, M. Cristina Pattuelli · Proceedings of the American Society for Information Science and Technology · 2005
In this paper, we present a reference model consisting of four different yet equally important perspectives for metadata requirements analysis. We argue that metadata development should start with a systematic requirements analysis as part of a systems engineering approach. This work has been developed as part of the GovStat project, an NSF-funded project designed to explore how US Federal statistical data can be made more understandable and usable by non-expert users. Developing a metadata architecture has been an important part of the project because we believe that applications to support finding and understanding materials from this unique data collection require a properly specified set of metadata. Since we were not able to find an existing metadata schema that meets our needs, we delved into a development of a new schema that was driven by the goals of the project, and by identifying crucial metadata elements through a series of metadata user studies (Denn, Haas, & Hert, 2003). In our effort to evaluate existing standards and define a new metadata schema, we realized that there is a lack of methodological principles and no established framework to inform the process. We believe that such a framework has emerged from our experience (Hert et al., in preparation) and is represented by the four-perspective model in Figure 1. Four Perspectives As for methodological principles, we propose to adopt a systems engineering approach. We believe this four-perspective model combined with systems engineering methodologies to be broadly applicable to metadata development initiatives. As in our case, there are often cases in which a unique set of requirements in the context of a specific project should be accommodated in a metadata architecture. Common approaches for this problem include adopting an existing metadata standard meeting the requirements, extending an existing standard, developing an application profile, or defining a new schema. Deciding upon a way to go is in and of itself challenging given the number of legitimate approaches and the abundance of metadata standards with varying degrees of complexity and diversity of scope and purpose. More importantly, whatever approach is taken, when we work on a metadata architecture in a specific context, multiple dimensions of requirements and constraints are involved. From what we have learned from our experience, we believe that the key to success is how thoroughly and clearly the requirements are defined. In fact, many metadata studies discuss specific requirements they are dealing with in terms of the characteristics of subject domain or resource format (Zeng, 1999; Boulos, 2004; Robson, 2001; Bird & Simons, 2003), the various needs of stakeholders (Sutton, 1999; Attig, Copeland, & Pelikan, 2004; Hert, Denn, & Haas, 2004), or the applications that will be built on top of the metadata (DuCharme, 2003). However, incorporation of these requirements is too often undertaken in an ad hoc fashion. Even though there are common components such as the consideration of the needs of the user groups, there seems to be no broadly adopted way to address them in the metadata development process. What we think is missing here is a well-grounded systematic approach to defining requirements and implementing them in a metadata architecture. We believe that introducing principles and practices from the field of Systems Engineering could be beneficial to address this problem. The concepts of systems engineering have evolved to cope with the complexity of development environments. The IEEE-Std 1220–1994 (1995) provides a formal definition: “Systems engineering (SE) is an interdisciplinary approach to derive, evolve, and verify a life cycle balanced system solution that satisfies customer expectations and meets public acceptability.” Systems engineering is about defining what should be done, and integrating and verifying what has been done throughout the life cycle of the development (Stevens, 1998). Its purpose is to create a viable (in the sense that it fulfills all the critical requirements), and effective (in the sense that it is achieved with minimum costs) solutions to complex problems. Not surprisingly, in systems engineering a clear emphasis is placed on the front end of the development life cycle, the analysis of requirements, because without properly defined and specifically documented requirements the goal is difficult to achieve. The importance of the requirements has led to the evolution of Requirements Engineering, which defines the systematic process of developing requirements through a variety of techniques and tools (Thayer & Dorfman, 1997). The issue of metadata requirements described above can be seen as a special case of the problems that have been well recognized and systematically approached in Systems Engineering and Requirements Engineering. Analysis first requires a frame of reference. As the first step toward what we are suggesting in this paper, we present the four-perspective model as a framework for identifying and classifying metadata requirements. It establishes organizing dimensions for the analysis. It also provides a high level abstraction upon which further specifications at various levels can be based. As shown in Figure 1, this model consists of four perspectives: content, user, organizational, and technical. These four perspectives are aligned on two axes: internal to the data producer vs. external to the data producer, and task-oriented vs. administration-oriented. Here we provide a brief description of the perspectives along with examples from our work. A content perspective describes the information space and its important features. Statistical data is presented in a variety of formats, primarily the statistical table. One of our stated goals is to be able to provide sub-document level access to tables, which requires developing metadata elements in such a way that the relationships between the components of a table can be preserved without the entire table itself needing to be preserved. A user perspective incorporates knowledge of the tasks that metadata needs to enable. For the purpose the GovStat project, the users we are primarily concerned with are non-experts. However, there are a wide range of users of Federal statistics, including educators, researchers, journalists, policy-makers, industry analysts, and the like. Each of these kinds of users has a different set of tasks that they might wish to undertake, which may require different metadata elements. An organizational perspective brings knowledge of the constraints and organizational aspects of information providers. There are whole systems of processes designed for administering surveys, collecting results, and performing the analyses that result in publicly available statistical information. Even a primarily user-oriented metadata schema must recognize the important organizational components that need to be included. A technical perspective addresses the use of appropriate standards and technologies, relationships to other components of the information and system architecture, and issues like extensibility, scalability, and interoperability. The US government statistical community has not adopted any standards for statistical metadata. In addition, standards relevant to our project are either very broad (ISO/IEC 11179) or focused narrowly on access to entire data sets from archives (DDI). These four perspectives affect the entire life span of metadata development and deployment. We believe it is important not only to analyze and elicit requirements from all of the four perspectives, but also to make it clear which requirement belongs to which perspective. We recognize that these perspectives are inter-related in many ways and thus it is not always possible to draw a line; however, trying to organize requirements and constraints from each and every perspective will give us a more complete understanding of multi-dimensional, multi-faceted problem situations and enable us to come up with a better strategy to coordinate different, even conflicting requirements. In current metadata development efforts, often some of these perspectives are missing in the discussion, and when these perspectives are addressed, they are not always clearly differentiated. We believe it is in part because there is no reference model that depicts all of these perspectives and their inter-relationships. General core requirements that should be met regardless of the domain are presented in the innermost layer. Domain-specific requirements stemming from the unique characteristics of the domain are modeled as the middle layer. Application-specific requirements relating to the expected use in the context of a specific local application are placed in the outer layer. Layered model of four perspectives For example, supporting sub-document access can be a general core requirement. Capturing information at different granularity levels in a table can be considered a domain-specific requirement. If the metadata needs to support a functionality such as manipulation of a table upon users' request, that would add some application-specific requirements such as providing elements that express what units the values in a table are reported in, what kinds of manipulations (such as summation, averaging, etc.) are appropriate for those units, etc. As the example above illustrates, the level of specificity increases when we move from the center to the outer parts of Figure 2, and in a sense we break down higher-level requirements into lower-level requirements. This would give us another advantage in the design process. When there are existing proposals or standards, we can evaluate each of them in terms of where they fit in the continuum of specificity and decide to adopt one or more standards (in the form of an application profile or an extension to existing solutions) depending on the desired level of specificity. With this framework for the analysis in place, the next step for Requirements Engineering would be to identify appropriate techniques and tools for gathering data, interpreting and analyzing the data, and defining and documenting requirements based on the analysis. For example, various techniques such as user interviewing, group session, etc. can be used for the data gathering process (Sommerville, 2001). Different techniques may be required for eliciting requirements for different perspectives. Again, the four-perspective model can guide the selection of a set of techniques suitable for metadata development. to analyze and document requirements / constraints. to assess what has been done and detect what is missing in every phase of the metadata development process to support harmonization of different or even conflicting requirements during the process to be used as a common basis for communication between stakeholders all the way from the requirements definition to the final adoption of the schema We believe that this approach will eventually enhance the adoptability of the metadata developed. Since this four-perspective model emerged during our metadata development process, we had already concluded the requirements analysis phase. Our plan is to use this model as a tool to evaluate the metadata elements we have created for the statistical domain. We believe this model has universal applicability in metadata development activities and that it can be helpful to information professionals and practitioners involved in any phase of metadata development to minimize risk of failure and conserve efforts. This work is supported by the National Science Foundation (NSF) under Grant NSF EIA 0131824.