Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2021 Jan 2;22(1):1.
doi: 10.1186/s12859-020-03881-z.

Propedia: a database for protein-peptide identification based on a hybrid clustering algorithm

Affiliations

Propedia: a database for protein-peptide identification based on a hybrid clustering algorithm

Pedro M Martins et al. BMC Bioinformatics. .

Abstract

Background: Protein-peptide interactions play a fundamental role in a wide variety of biological processes, such as cell signaling, regulatory networks, immune responses, and enzyme inhibition. Peptides are characterized by low toxicity and small interface areas; therefore, they are good targets for therapeutic strategies, rational drug planning and protein inhibition. Approximately 10% of the ethical pharmaceutical market is protein/peptide-based. Furthermore, it is estimated that 40% of protein interactions are mediated by peptides. Despite the fast increase in the volume of biological data, particularly on sequences and structures, there remains a lack of broad and comprehensive protein-peptide databases and tools that allow the retrieval, characterization and understanding of protein-peptide recognition and consequently support peptide design.

Results: We introduce Propedia, a comprehensive and up-to-date database with a web interface that permits clustering, searching and visualizing of protein-peptide complexes according to varied criteria. Propedia comprises over 19,000 high-resolution structures from the Protein Data Bank including structural and sequence information from protein-peptide complexes. The main advantage of Propedia over other peptide databases is that it allows a more comprehensive analysis of similarity and redundancy. It was constructed based on a hybrid clustering algorithm that compares and groups peptides by sequences, interface structures and binding sites. Propedia is available through a graphical, user-friendly and functional interface where users can retrieve, and analyze complexes and download each search data set. We performed case studies and verified that the utility of Propedia scores to rank promissing interacting peptides. In a study involving predicting peptides to inhibit SARS-CoV-2 main protease, we showed that Propedia scores related to similarity between different peptide complexes with SARS-CoV-2 main protease are in agreement with molecular dynamics free energy calculation.

Conclusions: Propedia is a database and tool to support structure-based rational design of peptides for special purposes. Protein-peptide interactions can be useful to predict, classifying and scoring complexes or for designing new molecules as well. Propedia is up-to-date as a ready-to-use webserver with a friendly and resourceful interface and is available at: https://bioinfo.dcc.ufmg.br/propedia.

Keywords: Clustering; Database; Peptides; Protein design; Protein structure; Protein–peptide complexes; Webserver.

PubMed Disclaimer

Conflict of interest statement

The authors declare that they have no competing interests.

Figures

Fig. 1
Fig. 1
Propedia database schema, presenting the tables, fields and relationships. The complex table (white) is the core of the database and interconnects all the data; pdb entities (blue) including group, pdb_groups and pdb tables; peptide/receptor and organism tables (yellow); cluster tables (green); and alignment tables (orange)
Fig. 2
Fig. 2
a Propedia scheme. The user accesses Propedia through a browser. Propedia presents each protein–peptide as a complex. Each complex can be associated with a cluster based on sequence, interface or binding site. b Propedia interface. Three-dimensional structure visualization of a complex. Protein is shown as a cartoon (alpha-helix in magenta and beta-strands in orange). The peptide is shown as a cartoon with cyan sticks. Complex information includes receptor features, peptide features, clustering classification and similar complexes. c–e Sequence, interface and binding site, cluster pages. Sequence cluster containing the sequence WebLogo (consensus) and main sequence. Each cluster page has a distribution chart (boxplot), used to filter complexes, according to the attributes used for clustering: sequence identity, iRMSD and alignment score
Fig. 3
Fig. 3
a Structural alignment between 2JF9 and 4IV2. The protein residues were conserved, but the peptide residues were not. b Estrogen receptor alpha LBD in complex with a tamoxifen-specific peptide antagonist (PDB id: 2jf9; peptide chain: Q; protein chain: B). c Estrogen receptor alpha ligand-binding domain in complex with dynamic way-derivative (PDB id: 4IV2; peptide chain: C; protein chain: A)
Fig. 4
Fig. 4
MEROPS specificity matrix in shades of blue and residues from Propedia suggested peptides highlighted in yellow
Fig. 5
Fig. 5
a PDB ID: 1lvb; peptide: chain D; protein: chain B; Rosetta score: − 538.306; Distance: 3.5 b PDB ID: 5om5; peptide: chain B; protein: chain A; Rosetta score: − 538.985; Distance: 3.7 c PDB ID: 1lvm; peptide: chain C; protein: chain A; Rosetta score: − 528.398; Distance: 5.5 d the whole set of evaluated peptides
Fig. 6
Fig. 6
Correlation of MetaD ΔGbind with site RMSD (left) and alignment score (right) from the Sars-Cov-2 MPro with peptide complexes from the PDB id: 2q6g (chain C), 1uk4 (chain H), 1lvm (chain D), and 1lvb (chain D)
Fig. 7
Fig. 7
AG’s Protease model, in gray, coupled with peptides 3qgn-A (a) and 4dii-L (b). The distance between the SER143 residue from the S1 site in the protease to the cysteine residues in the peptides are 3.9 Å and 4.4 Å respectively
Fig. 8
Fig. 8
AG’s Protease model, in gray, coupled with the 4 top scored poses of peptides 6rw2-B (a), 3kn2-B (b) and 2obq-B (c). Residues in red represent the catalytic residues from the catalytic triad

References

    1. Neduva V, Linding R, Su-Angrand I, Stark A, De Masi F, Gibson TJ, Lewis J, Serrano L, Russell RB. Systematic discovery of new recognition peptides mediating protein interaction networks. PLoS Biol. 2005;3(12):e405. doi: 10.1371/journal.pbio.0030405. - DOI - PMC - PubMed
    1. Liu D, Angelova A, Liu J, Garamus VM, Angelov B, Zhang X, Li Y, Feger G, Li N, Zou A. Self-assembly of mitochondria-specific peptide amphiphiles amplifying lung cancer cell death through targeting the vdac1-hexokinase-ii complex. J Mater Chem B. 2019;7(30):4706–4716. doi: 10.1039/C9TB00629J. - DOI - PubMed
    1. Lau JL, Dunn MK. Therapeutic peptides: historical perspectives, current development trends, and future directions. Bioorganic Med Chem. 2018;26(10):2700–2707. doi: 10.1016/j.bmc.2017.06.052. - DOI - PubMed
    1. Angelova A, Drechsler M, Garamus VM, Angelov B. Pep-lipid cubosomes and vesicles compartmentalized by micelles from self-assembly of multiple neuroprotective building blocks including a large peptide hormone pacap-dha. ChemNanoMat. 2019;5(11):1381–1389. doi: 10.1002/cnma.201900468. - DOI
    1. Lee AC-L, Harris JL, Khanna KK, Hong J-H. A comprehensive review on current advances in peptide drug development and design. Int J Mol Sci. 2019;20(10):2383. doi: 10.3390/ijms20102383. - DOI - PMC - PubMed