bio.rodeo
ModelsOrganizationsProvidersLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsProvidersLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
Protein

HelixFold-Multistate

Baidu PaddleHelix / Zonsen PepLib Biotech

GPCR-peptide complex structure prediction conditioned on active or inactive receptor states, used to rank designed peptide agonists and antagonists.

Released: November 2024

HelixFold-Multistate (HF-Multistate) is a structure prediction model for G protein-coupled receptor (GPCR)-peptide complexes that conditions its prediction on the receptor's functional state. GPCRs interconvert between active and inactive conformations, and which state a peptide stabilizes determines whether it behaves as an agonist or an antagonist. General-purpose complex predictors ignore this distinction: they return a single pose without regard to activation state, which makes their confidence scores a poor filter when the goal is a peptide with a specific pharmacology rather than mere binding.

The model was built by Baidu's PaddleHelix team with wet-lab partner Zonsen PepLib Biotech, and released as a bioRxiv preprint in November 2024 before publication in the Journal of Chemical Information and Modeling in October 2025. It is a fine-tune of HelixFold-Multimer, the team's PaddlePaddle reimplementation of AlphaFold-Multimer, trained further on GPCR-peptide and protein-peptide complexes and given state-conditioned template features in the manner of AlphaFold-Multistate. The result combines two capabilities that previously lived in separate models: AlphaFold-Multistate resolves activation states of monomeric GPCRs but models peptide interfaces poorly, while AlphaFold-Multimer models interfaces well but cannot be told which state to produce.

HF-Multistate serves as the scoring and filtering stage of a state-specific peptide design pipeline. Candidate sequences are generated elsewhere; the model predicts each candidate's complex with the target receptor in a specified state, discards complexes inconsistent with that state, and ranks the survivors by structural confidence. It is applied across distinct receptor targets without retraining, alongside GPCR-focused predictors such as RareFoldGPCR.

#Key Features

  • State conditioning: The receptor's active or inactive conformation is supplied as an input constraint through class- and state-matched structural templates, so agonist and antagonist candidates are evaluated against the conformation each is meant to stabilize.
  • GPCR-specific fine-tuning: Training on receptor-peptide complexes rather than generic protein interfaces measurably improves interface accuracy on held-out GPCR families.
  • Confidence scores that track affinity: Peptide pLDDT and interface pAE correlate more strongly with measured binding affinity than the equivalent AlphaFold-Multistate scores, making them usable as a screening statistic.
  • Transmembrane-domain fidelity: Accuracy is evaluated on TM3 and TM6 for class A receptors and TM3, TM6, and TM7 for class B, the helices whose displacement defines activation.
  • Efficient inference: A single model run with five random seeds reaches accuracy competitive with the five-model, five-seed ensembles used by the AlphaFold baselines.

#Technical Details

HF-Multistate inherits the Evoformer-plus-structure-module architecture of the AlphaFold-Multimer lineage. Fine-tuning used over 7,000 complexes drawn from Propedia v1.0 and GPCRdb with a June 30, 2021 cutoff, deduplicated across the two sources. Evaluation used 45 high-quality GPCR-peptide structures released after that cutoff (30 class A, 15 class B), selected so that no GPCR type overlaps the training set. On this benchmark HF-Multistate leads on DockQ and interface RMSD against AlphaFold-Multimer, HelixFold-Multimer, and AlphaFold-Multistate, and achieves the lowest RMSD across every key transmembrane helix. On a CXCR4-peptide affinity set, filtering candidates by confidence relative to a reference peptide yields a 42% hit rate (11/26) versus 33% for AlphaFold-Multistate; adding the requirement that confidence favor the correct receptor state raises it to 50% (11/22) versus 35%, against a 30% random baseline, where a hit is sub-micromolar affinity.

#Applications

In the published pipeline, FoldSeek retrieves peptide backbones from structurally similar complexes, an inverse folding model generates candidate sequences for those backbones, HF-Multistate predicts and filters the resulting complexes, and coarse MM-GBSA scoring selects a final shortlist for synthesis and calcium mobilization assays. Applied to three receptors, this produced APJR agonists with EC50 values of 4.2 nM and 5.0 nM, a GLP-1R antagonist at IC50 874 nM that outperforms the clinical-stage antagonist avexitide (1913 nM in the same assay), and GHSR agonists reaching 100 nM alongside antagonists in the 1.1–3.4 μM range. The workflow suits peptide therapeutic programs against receptors where the desired pharmacology, not just affinity, is the design objective.

#Impact

The work makes conformational state a first-class variable in structure-based peptide design and demonstrates that a fine-tuned complex predictor's confidence scores can prioritize candidates well enough to reach single-digit nanomolar potency from a few dozen syntheses. Its limitations are stated plainly by the authors: antagonist design lags agonist design, partly because antagonist peptide structures are scarce in the training data, and the GHSR campaign showed that inverse folding over a narrow set of seed backbones collapses sequence diversity. Neither model weights nor inference code have been released, so the model is not independently runnable, and the reported gains rest on one benchmark set and three targets.

Citations

DOI: 10.1021/acs.jcim.5c00884

Preprint

DOI: 10.1101/2024.11.27.625792

Recent citations

Papers that recently cited this model.

Not enough citation data yet.

Top citations

The most-cited papers that cite this model.

Not enough citation data yet.

Where to run HelixFold-Multistate

Providers that host HelixFold-Multistate for inference, fine-tuning, or weight download.

No providers recorded yet. Browse all providers

Fields of citing research

Not enough data

Openness

bio.rodeo opennessClosed · low usability and reproducibility
12Closed
Usability — can I run it?9
Reproducibility — can I retrain it?16

Tags

peptide_designpeptidesstructure_predictiontransfer_learningtransformer

Resources

Research PaperResearch Paper