Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2025 Jan;44(1):e202400186.
doi: 10.1002/minf.202400186. Epub 2024 Oct 10.

Navigating a 1E+60 Chemical Space of Peptide/Peptoid Oligomers

Affiliations

Navigating a 1E+60 Chemical Space of Peptide/Peptoid Oligomers

Markus Orsi et al. Mol Inform. 2025 Jan.

Abstract

Herein we report a virtual library of 1E+60 members, a common estimate for the total size of the drug-like chemical space. The library is obtained from 100 commercially available peptide and peptoid building blocks assembled into linear or cyclic oligomers of up to 30 units, forming molecules within the size range of peptide drugs and potentially accessible by solid-phase synthesis. We demonstrate ligand-based virtual screening (LBVS) using the peptide design genetic algorithm (PDGA), which evolves a population of 50 members to resemble a given target molecule using molecular fingerprint similarity as fitness function. Target molecules are reached in less than 10,000 generations. Like in many journeys, the value of the chemical space journey using PDGA lies not in reaching the target but in the journey itself, here by encountering non-obvious analogs. We also show that PDGA can be used to generate median molecules and analogs of non-peptide target molecules.

Keywords: chemical space; cheminformatics; genetic algorithm; therapeutic peptides.

PubMed Disclaimer

Conflict of interest statement

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Figures

Figure 1
Figure 1
Design of PDGA. PDGA uses a list of input building blocks to generate a set of random linear sequences. The sequences are encoded using either the MAP4 C or MXFP fingerprints. The fingerprints are used to determine the fitness of the sequences by calculating the distance towards a specified query molecule. Sequences with distances below a set threshold are stored in an analogs database. The 15 fittest sequences undergo rounds of mutations and crossovers in which building blocks and topology are changed to add 35 new sequences to the population. This process iterates until either the query is found or the PDGA reaches 10,000 generations.
Figure 2
Figure 2
Structures of the selected queries for the PDGA runs using the MAP4 C similarity as fitness function. The linear sequences (4, 5 and 6) are written with standard one‐letter code for amino acids, with free N‐terminus marked as “H−“ and C‐terminus in acid form “−OH” or amide form “−NH2”.
Figure 3
Figure 3
Analysis of three parallel PDGA runs starting from 50 random sequences towards selected queries. Top plots show the overall best score throughout the trajectory; the bottom plots show the cumulative number of unique new molecules generated throughout the trajectory for a) polymyxin B2, b) EB9, and c) cathelicidin BF. The best score refers to the MAP4 C dJ of the closest structure generated up to that generation relative to the target.
Figure 4
Figure 4
Analysis of polymyxin B2 runs starting from 50 random linear sequences (Run 1–3) or from polymyxin B2 without stopping condition (Self). a) Heatmap indicating the number of generated compounds with MAP4 C dJ <0.5 to polymyxin B2 for each trajectory, along with the number of overlapping compounds. b) Bar plot showing the mean and standard deviation of the dJ calculated using MAP4 C fingerprints for generated compounds with dJ <0.5 to polymyxin B2. c) Bar plot showing the mean and standard deviation of the Levenshtein distance (dL ; proxy for number of mutations) to polymyxin B2 for generated compounds with dJ <0.5 to polymyxin B2. d) Structure of a selected polymyxin B2 analog featuring a high dL and low dJ (7) and the closest analog generated in the failed run (8). e) TMAP displaying the generated compounds in a 2D space. Interactive TMAP: https://tm.gdb.tools/map4/10E60/polymyxin_randself_tmap.html.
Figure 5
Figure 5
Visualization of traversal trajectories and median molecules between polymyxin B2 and gramicidin S. a) Jaccard distance of molecules selected from the different trajectories towards polymyxin B2 and gramicidin S. The trajectory from polymyxin B2 to gramicidin S is displayed in blue, the reverse trajectory is displayed in red, and the combined structure trajectory is displayed in yellow. b) MAP4 C TMAP of selected molecules colored by their trajectory of origin. The trajectories populate separate chemical subspaces. c) Structures of the two queries polymyxin B2 and gramicidin S and two selected molecules from the median trajectory (yellow). Interactive TMAP: https://tm.gdb.tools/map4/10E60/polymyxin_gramicidin_tmap.html.
Figure 6
Figure 6
Non‐peptide macrocycles, the overall best score throughout the trajectories and the corresponding best scoring MXFP analog from three combined runs for a) cyclosporin and b) valinomycin. The MXFP dCBD is reported for each analog. See also Figure S6 for further details.

References

    1. Lam K. S., Salmon S. E., Hersh E. M., et al., “A New Type of Synthetic Peptide Library for Identifying Ligand-Binding Activity”, Nature 354, no. 6348 (1991): 82–84, 10.1038/354082a0. - DOI - PubMed
    1. Houghten R. A., Pinilla C., Blondelle S. E., et al., “Generation and Use of Synthetic Peptide Combinatorial Libraries for Basic Research and Drug Discovery”, Nature 354, no. 6348 (1991): 84–86, 10.1038/354084a0. - DOI - PubMed
    1. Lam K. S., Lebl M., Krchňák V., “The ‘One-Bead-One-Compound’ Combinatorial Library Method”, Chemical Reviews 97, no. 2 (1997): 411–448, 10.1021/cr9600114. - DOI - PubMed
    1. Bohacek R. S., McMartin C., Guida W. C., “The Art and Practice of Structure-Based Drug Design: A Molecular Modeling Perspective”, Medicinal Research Reviews 16, no. 1 (1996): 3–50, https://doi.org/10.1002/%28SICI%291098-1128%28199601%2916:1<3::AID-MED1>3.0.CO;2-6. - PubMed
    1. Bleicher K. H., Bohm H. J., Muller K., et al., “Hit and Lead Generation: Beyond High-Throughput Screening”, Nature Reviews Drug Discovery 2, no. 5 (2003): 369–378, 10.1038/nrd1086. - DOI - PubMed