The pdb2sql Python Package: Parsing, Manipulation and Analysis of PDB Files Using SQL Queries
Nicolas Renaud, Cunliang Geng · The Journal of Open Source Software · 2020
The analysis of biomolecular structures is a crucial task for a wide range of applications ranging from drug design to protein engineering.The Protein Data Bank (PDB) file format (Burley et al., 2019) is the most popular format to describe biomolecular structures such as proteins and nucleic acids.In this text-based format, each line represents a given atom and entails its main properties such as atom name and identifier, residue name and identifier, chain identifier, coordinates, etc.Several solutions have been developed to parse PDB files into dedicated objects that facilitate the analysis and manipulation of biomolecular structures.This is, for example, the case for the BioPython parser (Cock et al., 2009,@biopdb) that loads PDB files into a nested dictionary, the structure of which mimics the hierarchical nature of the biomolecular structure.Selecting a given sub-part of the biomolecule can then be done by going through the dictionary and selecting the required atoms.Other packages, such as ProDy (Bakan, Meireles, & Bahar, 2011), BioJava (Lafita, 2019), MMTK (Hinsen, 2000) and MDAnalysis (Gowers et al., 2016) to cite a few, also offer solutions to parse PDB files.However, these parsers are embedded in large codebases that are sometimes difficult to integrate with new applications and are often geared toward the analysis of molecular dynamics simulations.Lightweight applications such as pdb-tools (Rodrigues, Teixeira, Trellet, & Bonvin, 2018) lack the capabilities to manipulate coordinates.We present here the Python package pdb2sql, which loads individual PDB files into a relational database.Among different solutions, the Structured Query Language (SQL) is a very popular solution to query a given database.However SQL queries are complex and domain scientists such as bioinformaticians are usually not familiar with them.This represents an important barrier to the adoption of SQL technology in bioinformatics.pdb2sql exposes complex SQL queries through simple Python methods that are intuitive for end users.As such, our package leverages the power of SQL queries and removes the barrier that SQL complexity represents.In addition, several advanced modules have also been built, for example, to rotate or translate biomolecular structures, to characterize interface contacts, and to measure structure similarity between two protein complexes.Additional modules can easily be developed following the same scheme.As a consequence, pdb2sql is a lightweight and versatile PDB tool that is easy to extend and to integrate with new applications.