Joshua Almonte, PhD
Cornell University
Generate Broadly, Screen Selectively: In-Context Enrichment for Peptide Hit Identification
Early-stage peptide discovery requires searching large sequence spaces while reserving expensive structure-based evaluation for the most promising candidates. We are developing a computational screening approach that combines BoltzGen with Multi-Peptide Example Prompting (MPEP), an in-context learning method for protein language models capable of scoring 10^7 peptide sequences/second. In our workflow, BoltzGen’s comparatively inexpensive design and inverse-folding stages generate an expanded candidate library, approximately tenfold larger than a conventional campaign. MPEP then serves as an orthogonal pre-screen to prioritize candidates before full structure prediction and filtering, increasing search breadth for a fixed, total computational budget. In an initial proof-of-concept using the MHC-II protein HLA-DR1 as a target, MPEP scores showed consistent positive correlations with 5/7 downstream BoltzGen prediction and ranking metrics (Spearman correlation=0.3–0.4). These results support sequence-based in-context scoring as a lightweight prioritization step that may enrich high-quality peptide candidates entering hit identification and improve the efficiency of early-stage peptide discovery.
Early-stage peptide discovery requires searching large sequence spaces while reserving expensive structure-based evaluation for the most promising candidates. We are developing a computational screening approach that combines BoltzGen with Multi-Peptide Example Prompting (MPEP), an in-context learning method for protein language models capable of scoring 10^7 peptide sequences/second. In our workflow, BoltzGen’s comparatively inexpensive design and inverse-folding stages generate an expanded candidate library, approximately tenfold larger than a conventional campaign. MPEP then serves as an orthogonal pre-screen to prioritize candidates before full structure prediction and filtering, increasing search breadth for a fixed, total computational budget. In an initial proof-of-concept using the MHC-II protein HLA-DR1 as a target, MPEP scores showed consistent positive correlations with 5/7 downstream BoltzGen prediction and ranking metrics (Spearman correlation=0.3–0.4). These results support sequence-based in-context scoring as a lightweight prioritization step that may enrich high-quality peptide candidates entering hit identification and improve the efficiency of early-stage peptide discovery.
