PROMIS: An XML‐based metadata framework for proteomics
W. John MacMullen · Proceedings of the American Society for Information Science and Technology · 2003
Scientists who study proteins generate and manipulate large quantities of primary data in the course of determining the proteins' composition, three-dimensional structure, and function. They use a wide variety of physical instruments and computational methods to perform these analyses. The resulting data sets and associated metadata may or may not be comparable, depending upon the methods used to derive and capture them. To date, no standard way of capturing, managing, and sharing these data and metadata has been developed. As a result, it is difficult for researchers to integrate homogeneous data sets from different experiments, collect metadata consistently across experiments, share data sets with colleagues, and perform analysis across (and create relationships among) heterogeneous data sets (e.g., analyzing protein sequence data with gene expression data). The PROMIS (PROteomics Metadata Interchange Schema) project described in this presentation is a proof-of-concept prototype of an open metadata standard for compositional proteomics. Proteomics is the systematic study of the expression, composition, location, quantity, structure, function, and interaction of proteins in organisms, and the variation of those factors between normal and perturbed proteins. As a proof-of-concept, this project is focused on compositional aspects, which refers to the primary structure of proteins; that is, the chemical constituents of translated proteins and the determination of their amino acid seq-uences. This project encompasses general classes of data and information that describe entities such as experimental design, equipment, samples, controls, and measurements. This work advances the framework created in MacMullen, et al. (2002) by extending the defined element and attribute sets, formalizing them into a semantic network, and developing an XML-based DTD for use in sharing data sets and their related metadata. The end goal is an extensible, instrument-, system-, and vendor-independent standard for describing and annotating compositional proteomics data and metadata. The framework is being developed for testing by UNC's Center for Bioinformatics, Proteomics Core Facility, and a proteomics-focused lab in the Department of Microbiology and Immunology. The architecture will be based on open-source software components, and will provide form-based interfaces to enable author-generated metadata for attributes that are not automatically generated. Information on proteomics objectives, technologies, protocols, and workflows was gathered via literature reviews and interviews with UNC Proteomics Core facility directors, personnel in the Center for Bioinformatics, and investigators employing proteomics methods in their research. The proteomics information described above was organized and developed into a semantic network, of which Figure 1 is a partial example. The focal point of the network is a particular biological sample, which has a variety of attributes (including its biological and acquisition sources and the associated investigator). PROMIS semantic network (partial) Two types of processes are performed within the proteomics workflow at the highest level: physical sample processing (e.g., determining its chemical constituents), and analysis of the data resulting from physical processing (e.g., determining the similarity of an unknown sample's sequence to known sequences). Each of these processes has associated hardware and software devices, with protocols and parameters that vary according to each individual experiment. An RDF/XML document type definition (DTD) is under development for PROMIS. Figure 2 provides an example of the Physical Processing element, in this case illustrating a type of process (sequencing), which uses a type of device (mass spectrometer) and a specific protocol. PROMIS RDF/XML DTD example RDF (Resource Description Framework) (W3C, 2003) was selected because of its ability to represent relationships using a formal semantics to enable inferencing, and its ability to describe resources that exist in formats other than web pages. The former is important for automating discovery of implicit relationships among biological entities, while the latter is important because most data sets are not stored online as static web pages. RDF is also used by other leading bioinformatics representation schemes, such as the Gene Ontology resource (GO, 2003). Proteomics is a new discipline whose technologies are evolving rapidly, so one design goal for PROMIS has been to generalize entities, processes, and workflows to remain independent of technologies. In addition, following Holsapple & Joshi (2002), a related goal is to provide a means for describing data, not for prescribing methodologies. Additional elements and attributes can be added as needed to enrich the level of detail. In practice, the PROMIS framework will be implemented at multiple levels that require additional situation-specific components. An obvious use is as a common format for data set interchange. But that assumes data sets have been structured according to the PROMIS semantic network and DTD. The nature of proteomics workflow at UNC - a centralized core facility to which labs and investigators submit samples for processing - suggests that author-generated metadata at the time of sample submission could be an option. A later phase of this project will create a web-based interface with controlled vocabularies for attributes and values, and functionality that allows investigators to create profiles to quickly load repetitive samples. Crosswalks will also be needed to map fields from existing sample tracking databases and laboratory information management systems (LIMS) to PROMIS for data set exchange. One planned deliverable is the definition of a required minimal element set to enable other labs to test PROMIS in their own settings. Additionally, we want to ensure leverage of (or, minimally, consistency with) protein-specific elements of emerging standards, such as the Biopolymer Markup Language (BIOML, 2003) for those experiment-independent parts of PROMIS that describe protein composition (e.g., amino acids), structure, or function. Since investigators are interested in the relationships between proteins and gene expression, consistency with the MIAME (Minimum information about a microarray experiment) draft standard (MIAME, 2003) is also desired. This work was funded in part by NLM training grant LM07071 to the Department of Biomedical Engineering, School of Medicine, and a fellowship from the Program in Bioinformatics and Computational Biology in the Carolina Center for Genome Sciences, University of North Carolina, Chapel Hill. The author would also like to thank the coauthors of the prior work this presentation extends, and the investigators and research staff who provided information and feedback throughout the requirements definition and design processes.