Metadata foundations for the life cycle management of software systems
David Hyland-Wood · The University of Queensland · 2008
Software maintenance is, often by far, the largest portion of the software lifecycle in terms ofnboth cost and time. Yet, in spite of thirty years of study of the mechanisms and attributes ofnmaintenance activities, there exist a number of significant open problems in the field: Software stillnbecomes unmaintainable with time.nSoftware maintenance failures result in significant economic costs because unmaintainablensystems generally require wholesale replacement. Maintenance failures occur primarily becausensoftware systems and information about them diverge quickly in time. This divergence is typicallyna consequence of the loss of coupling between software components and system metadata thatneventually results in an inability to understand or safely modify a system. Accurate documentationnof software systems, long the proposed solution for maintenance problems, is rarely achieved andneven more rarely maintained. Inaccurate documentation may exacerbate maintenance costs when itnmisdirects the understanding of maintainers.nThis thesis describes an approach for increasing the maintainability of software systems via thenapplication and maintenance of structured metadata at the source code and project descriptionnlevels. The application of specific metadata approaches to software maintenance issues isnsuggested for the reduction of economic costs, the facilitation of defect reduction and the assistancenof developers in finding code-level relationships. The vast majority of system metadata (such asnthat describing code structure, encoded relationships, metrics and tests) required for maintainabilitynis shown to be capable of automatic generation from existing sources, thus reducing needs fornhuman input and likelihood of inaccuracy. Suggestions are made for dealing with metadata relatingnto human intention, such as requirements, in a structured and consistent manner.nThe history of metadata is traced to illustrate the way metadata has been applied to virtual,nphysical and conceptual resources, including software. Best practice approaches for applyingnmetadata to virtual resources are discussed and compared to the ways in which metadata havehistorically been applied to software. This historical analysis is used to explain and justify anspecific methodological approach and place it in context.nTheories describing the evolution of software systems are in their infancy. In the absence of anclear understanding of how and why software systems evolve, a means to manage change isndesperately needed.nA methodology is proposed for capturing, describing and using system metadata, coupling itnwith information regarding software components, relating it to an ontology of software engineeringnconcepts and maintaining it over time. Unlike some previous attempts to address the loss ofncoupling between software systems and their metadata, the described methodology is based onnstandard data representations and may be applied to existing software systems. The methodology isnsupportive of distributed development teams and the use of third-party components by itsngrounding of terms and mechanisms in the World Wide Web.nScaling the methodology to the size of the Web required mechanisms to allow the Webrs URL nresolution process to be curated by metadata maintainers. Extensions to the Persistent URL conceptnwere made as part of this thesis to allow for curation of URL resolutions. A computationalnmechanism was defined to allow for the distinction of virtual, physical and conceptual resourcesnthat are identified by HTTP URLs. That mechanism was used to represent different types ofnresources in RDF using HTTP URLs while avoiding ambiguities.nThe maintenance of system metadata is shown to facilitate understanding and relationalnnavigation of software systems and thus to forestall maintenance failure. The efficacy of thisnapproach is demonstrated via a modeling of the methodology and two case studies.n