Exploring the Chemical Space of Antiparasitic Peptides and Discovery of New Promising Leads through a Novel Approach based on Network Science and Similarity Searching - Supporting Information
Sebastián Ayala‐Ruano, Yovani Marrero‐Ponce, Longendri Aguilera‐Mendoza, Noel Pérez, Guillermı́n Agüero-Chapin, Agostinho Antunes, Ana Cristina Aguilar · Zenodo (CERN European Organization for Nuclear Research) · 2021
Supporting Information for a draft of the article titled Exploring the Chemical Space of Antiparasitic Peptides and Discovery of New Promising Leads through a Novel Approach based on Network Science and Similarity Searching. We explain the contents of this material below: Supporting information 1 (SI1): FASTA files of 550, 415, and 405 antiparasitic peptides (APPs) to generate networks for this study. Also, this folder contains FASTA files of the five benchmarking datasets of APPs/non-APPs used to compare the performance of our multi-query similarity searching models (mQSSMs) between them and with the algorithms previously reported in the literature. Supporting information 2 (SI2): MS word file with tables of parameters of similarity threshold analysis for chemical space networks (CSNs) and half-space proximal networks (HSPNs) of APPs; common APPs in the top 50 of most central nodes from CSN and HSPN retrieved by weighted degree, hub-bridge, harmonic, and betweenness centrality measures; the number and community membership percentage of the most central APPs by weighted degree, hub-bridge, harmonic, and betweenness centrality measures; and query sets and the number of queries from the mQSSMs to find novel APPs for CSN, HSPN, and CSN-HSPN. Supporting information 3 (SI3): Graphml files of all the networks created in this study. Supporting information 4 (SI4): Excel file with normalized weighted degree, hub-bridge, harmonic, and betweenness centrality measures for CSN and HSPN. Supporting information 5 (SI5): FASTA files of the most central and non-redundant APPs by each type of network and centrality measures, or the query sets. Supporting information 6 (SI6): Excel files with results of mQSSMs (Output Predictions). Supporting information 7 (SI7): This folder has three kinds of files, SI7-A that contains three folders with original results for each mQSSM generated, as well an excel file with statistical parameters. SI7-B is an excel file with the performance parameter of the best 21 mQSSMs proposed, as well as the ranking of these models. Finally, SI7-C and SI7-D are pdf files with results of multiple comparisons of our mQSSMs and with literature algorithms, respectively. Supporting information 8 (SI8): FASTA files of 95 lead compounds and communities obtained from the CSN of these compounds. Supporting information 9 (SI9): PowerPoint file with alignments and sequence logos of 95 lead compounds and communities obtained from the CSN of these peptides.