Grassroots: An infrastructure for sharing services & data
Xingdong Bian, Simon Tyrrell, Robert Davey · 2019
Earlham Institute, Norwich Research Park, Norwich, NR4 7UZ, United Kingdom Integrative research requires extensive multi-level approaches to enrich and expose data and workflows so that informatics infrastructures can process them effectively. As part of the Designing Future Wheat (DFW) project, the Grassroots Infrastructure has been developed at the Earlham Institute (EI) to consolidate data and analyses, facilitating consistent approaches to generating, processing and disseminating public datasets in the plant sciences. It is also part of the Wheat Initiative Wheat Information System (WheatIS) project, formalising the infrastructure as the federated UK WheatIS node involving partners from the University of Bristol, the European Bioinformatics Institute, Rothamsted Research, and the John Innes Centre. Its lightweight reusable software stack comprises: an iRODS data management layer to provide structure to unstructured file systems and WebDAV APIs exposed via EIRods-DAV; interfaces to interact with local or cloud-based analysis platforms; an Apache web server layer to deliver content and provide access to public programmatic interfaces; services such as: BLAST searches on multiple databases across different sites, systems for storing and searching field trial experiments, a mapping tool showing pathogen samples with temporal and spatial data, and adding API layers to the SeedStor application by the Germplasm Resources Unit allowing it to make seamless connections to research objects stored on our new DFW data portal. To make the data as reusable as possible, all data is marked up using appropriate ontologies. The Grassroots Infrastructure can be run locally or packaged in virtual containers and deployed on a variety of hardware thus representing a decentralised system, allowing information generators to retain control over their resources but allowing interconnected resources to access each other consistently. We are currently working on various enhancements to allow for more functionality and user-friendliness. These include lightweight mechanisms to expose underlying grid architecture by extending our WebDAV solution still further to allow for metadata interaction, adopting standardised APIs such as the Breeding API (BrAPI) and schemas such as Frictionless Data, BioSchemas, etc. These will enable greater interoperability with a variety of existing services and data analysis platforms such as CyVerse and Galaxy. We will enable integration with CKAN, to share, preserve, cite, explore, and analyse research data, along with our own custom stools to manage and visualise trial data. Further data mining will be available with a cohesive context-aware search engine using a Lucene-based indexing tool, designed with wheat-based data in mind, across all of the connected data and services. It is fully open source and available on GitHub. More information can be found at https://grassroots.tools