Graphlet based network analysis of protein structures

Ioannis Filippis · Universitätsbibliothek der FU Berlin Hochschulschriftenstelle u. Dokumentenserver · 2012

Network analysis of protein structures has provided valuable insight into protein folding and function. However, the lack of a unifying view in network modelling and analysis of protein structures and the unexploited advances in network theory prompted me to address three important challenges: 1\. Rationalise the choice of network representation of protein structures. 2\. Propose a well fitting null model for protein structure networks. 3\. Develop a novel graph-based whole-residue empirical potential. Graphlets, a recently introduced and powerful concept in graph theory, are a fundamental aspect of this thesis. The topological similarity between protein structure networks or individual residues was assessed using graphlet-based methods in order to propose an optimised null model and develop a novel potential. Chapter 2 unifies the view of network representations by means of a controlled vocabulary and outlines the motivation behind the details of constructing such networks, and the popularity and optimality of the representations. In Chapter 3, an exhaustive set of 945 network representations is systematically analysed with respect to their similarity and fundamental network properties. The similarity between commonly used representations can be quite low and specific representations may exhibit high number of orphan residues and residues lying in ”separate” components. Additionally, proteins with different secondary structure topologies have to be treated with caution in any network analysis. This work allows for a rational selection of a network representation based on certain principles, popularity, optimality and desired network properties and on its similarity to successfully utilised representations. Chapter 4 shows that 3-dimensional geometric random graphs, that model spatial relationships between objects, provide the best fit to protein structure networks among several random graph models. The fit is overall better for a structurally diverse protein data set, various network representations and with respect to various topological properties. Geometric random graphs capture the network organisation better for larger proteins and proteins of low helical content and low thermostability. Choosing geometric random graphs as a null model results in the most specific identification of statistically significant subgraphs. In Chapter 5, a novel knowledge-based potential is developed by generalising the single-body contact-count potential to a whole-residue pure topological one. The proposed scoring function outperforms the contact-count potential. The improved performance is consistent across various methods of generating decoys with respect to most performance metrics and is more prominent for the most successful fragment-based methods. The potential is also on par with a traditional four-body potential and exhibits strong complementarities with it, highlighting the capacity for further improvement. Overall, this dissertation establishes the basis for the analysis of protein structures as networks and opens the door to new avenues in the quest for the perfect energy function.

Read the paper · More papers on PaperTik