Scalable, Interactive, and Reproducible Data Mining of 3D Macromolecular Structures
Shih-Cheng Huang, Yue Yu, Peter W. Rose · 2019
The Protein Data Bank (PDB) represents the core data resource for Structural Bioinformatics. The rapid growth of the PDB (> 150,000 structures) enables large-scale data mining, such as development of knowledge-based potentials, docking and scoring functions, and machine learning for protein structure and function prediction. We have developed efficient data representations ( MacroMolecular Transmission Format ) and a scalable framework to mine the PDB using state-of-the-art Big Data Technologies ( mmtf-pyspark ). We have deployed applications of this framework in Jupyter Notebooks that are hosted on free public servers, including mybinder.org and CyVerse.org, enabling researchers to publish documented workflows that are reproducible and that can be re-run, modified, or used as starting-points for new structural analyses. We present our approach of using Apache Spark and columnar data formats to scale structural analysis to enable the interactive exploration of the PDB archive, as well as scalable data integration. We demonstrate these capabilities by creating representative subsets of the PDB for machine learning applications and the mapping and visualization of post-translational modifications from proteomics experiments ( mmtf-proteomics ) and genomic variations to 3D structures ( mmtf-genomics ) in the context of protein-protein/nucleic acid/ligand/drug interactions. We also cover best practices of deploying these workflows on public servers to enable reproducibility and reuse .