On the interest of semi-supervised approaches with spacial dependence in structural alphabet encoding

Ikram Allam, Delphine Flatters, Leslie Regad, Anne‐Claude Camproux, Grégory Nuel · 2015

Structural alphabet (SA) has proved to be a powerful tool to compress the three dimensional protein conformations (3D) of proteins into a one-dimensional representation (1D) by encoding protein fragments into structural sequences or Structural Letter sequences (SLs) which can be analysed using standard sequence analysis Tools. SA has demonstrated its usefulness for protein analysis through many applications such as protein classification, structure alignment, structure fast comparison, extraction of functional motifs, etc. The development of a SA able to integrate the flexibility of the 3D structure of proteins and their cavity (pocket) capable of binding drug is a crucial challenge in the drug design and drug discovery domains. This is now possible thanks to the growing number of protein 3D structures identified in the 3D protein databank PDB (≈99,642 structures). In this work, we present a brief review of the various SA available to encode 3D PDB structures into SL. These various SA are based on mixture models, hidden Markov models (HMM) or classification tools. We then evaluate the interest of two fundamental concepts: a) unsupervised or semi-supervised training of SA; b) accounting or not for the spacial dependence between protein fragments. The results show that the integration of the semi-supervised approach with fragment spacial dependence, typically through HMM, can dramatically improve the performance of SA. It could be a very promising approach for drug design applications.

Read the paper · More papers on PaperTik