Abstract 4503: Analyzing and expanding the druggable proteome from predicted models

Phillip W Gingrich, Ansuman Biswas, Bissan Al‐Lazikani · Cancer Research · 2025

Abstract Introduction: The number of experimentally solved protein structures is ever-growing. These are vital for informing early drug discovery, yet only 40% of the human proteome is structurally characterized experimentally. Newmodelling tools such as AlphaFold are heralded to plug the gap in structural knowledge for the drug discoverer. Importantly, what part of the human proteome is available to drug discovery using the collective data? Here, we report analyses at the residue, protein, and cluster levels to understand how much of the proteome is structurally enabled for drug discovery, and we present the first comprehensive, reliable druggability map of the human proteome. Methods: We mapped experimental structures and AF models to canonical forms in the human proteome, mapping experimental coverage and pLDDT scores to each residue. As a proxy for useful domain-sized structures, the presence of at least 100 consecutive residues with experimental coverage or high prediction confidence was annotated per protein. MMSeqs2 clustering was used to explore proteins where AF struggled, both with and without experimental coverage in the PDB. Finally, protein pockets were found using SURFNET and classified by our draggability classifier. Results: The PDB covers 2.9 million residues across the proteome. AF affords 4.2 million additional confident predictions, yielding 68% of the proteome as structurally enabled at the residue level. 7100 proteins of those in the PDB have >= 100 consecutive residues with experimental coverage. AF provides 7600 more proteins with >= 100 consecutive confident residues, resulting in 72% of the proteome being structurally enabled with a domain-sized structure. AF makes low confidence predictions for the last 5700 proteins, most of which no experimental coverage. Likewise, low confidence predictions are made for proteins typically in complex with other macromolecules, despite 75% of them having at least 50% experimental coverage. In total, 40% of the proteome was classified as druggable using experimental structures or high confidence models, with as much as 60% being druggable with loosened confidence criteria. We provide a detailed map and guidelines on the utility of AF for drug discovery. Conclusion: Roughly 70% of the proteome has been made accessible and suitable for in silico drug discovery efforts based on experimental coverage and prediction uncertainty. This has expanded the druggable fraction of the proteome significantly over previous estimates. This is the first comprehensive analysis on the suitability of predicted models, highlighting where predicted models are practically useful to the drug discoverer. Citation Format: Phillip W. Gingrich, Ansuman Biswas, Bissan Al-Lazikani. Analyzing and expanding the druggable proteome from predicted models [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 4503.

Read the paper · More papers on PaperTik