Discovering 3D Protein Structures for Optimal Structure Alignment
Tomá Novosád, Václav Snáel, Ajith Abraham, Jack Y. Yang · 2013
Analyzing three-dimensional protein structures is a very important task in molecular biology. It has been proved that structurally similar proteins tend to have similar functions even if their amino acid sequences are not similar to one another. Thus, it is very important to find proteins with similar structures from the growing database to analyze protein functions. Currently there exist several protein databases publicly available online. These databases assemble various data about proteins, protein structures, protein functions, protein relationships, and other information. This chapter describes the procedure for building the matrix representing the vector model index file. It discusses the suffix tree data structure-its definition, construction algorithms, and main characteristics. The data for protein 3D structures indexing are retrieved from Protein Data Bank (PDB) database, which consists of proteins, nucleic acids, and complex assemblies. The chapter further explains the algorithm for measuring protein similarity on the basis of their tertiary structure.