Abstract 2043 Predicting Functions of Uncharacterized Proteins in Prokaryotes Using a Combination of Structural and Sequence-Based Approaches

Diya Mathai, Stefan Schulze · Journal of Biological Chemistry · 2025

Mass spectrometry-based proteomics regularly results in the identification of proteins that are differentially abundant in analyzed conditions, but whose functions are unknown so far.Especially in prokaryotes, these proteins of unknown function take up large parts of the proteome and are critical to deciphering the pathogenicity of bacteria like Pseudomonas aeruginosa.While sequence-based tools have been utilized in predicting protein functions for years, recent advances in structure prediction tools, such as AlphaFold, now allow for structure-based comparisons that provide additional insights into protein function.This study integrates sequence and structure-based computational tools, including EggNOG-mapper and Foldseek, to predict the functions of uncharacterized proteins.The prediction pipeline begins with FASTA sequences, which are used in EggNOG-mapper to provide functional annotations based on sequence homology.In addition, AlphaFold structure models are acquired from UniProt and used in Foldseek to identify structural homologs through structure-based comparisons.By combining results from sequence and structure-based searches, this pipeline improves the comprehensiveness of functional predictions.The pipeline was applied to quantitative proteomics datasets for P. aeruginosa, in which a large proportion of differentially abundant proteins remain to be characterized .In the P. aeruginosa reference proteome PAO1, comprising 5,586 protein-encoding genes, 66.2% are categorized as having unknown (548), hypothetical (2,045), or probable (1,103) functions.We analyzed a dataset where, following betalactam antibiotic exposure, 33% of the differentially abundant proteins were of unknown or hypothetical function.Applying our pipeline revealed functional insights into these proteins, providing further insights into their roles in the antibiotic resistance of P. aeruginosa.Our findings demonstrate that combining structure-and sequence-based approaches yields complementary insights.The developed pipeline can be easily extended to incorporate results from newly developed annotation tools, and is applicable to protein sequences from virtually any species.When coupled with quantitative proteomics data, this integrated approach enables a more comprehensive understanding of previously uncharacterized proteins and the cellular pathways they are involved in.

Read the paper · More papers on PaperTik