bio.rodeo
ModelsOrganizationsProvidersLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsProvidersLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
ProteinSmall molecule

HydrAffinity

Lanzhou University

Preparation-free protein-ligand binding affinity prediction from a protein sequence and a ligand SMILES, using a cascaded mixture-of-experts fusion.

Released: July 2026

Protein-ligand affinity prediction splits into two camps. Interaction-based scoring functions dock the ligand, build a 3D complex, and read off atom-level contacts — accurate, but the conformation-preparation step is the bottleneck that makes billion-compound screening impractical. Interaction-free methods skip docking entirely, taking only a receptor amino-acid sequence and a ligand SMILES string, and pay for that throughput with markedly lower accuracy. HydrAffinity, from Huiming Bao and Shouliang Dong at Lanzhou University, is an attempt to close that gap without reintroducing structure preparation.

The work begins with an audit rather than an architecture. The authors benchmark sequence-, graph-, and image-based pretrained encoders on the same affinity task under a common projector–transformer–predictor harness, producing a like-for-like comparison of encoder families that the field had lacked. SMILES-based transformers consistently beat graph and image encoders; among protein encoders, larger pretraining helps, but the best ESM-2 representation comes from an intermediate layer rather than the final output — evidence that fine-grained features carry more affinity signal than coarse abstractions.

The second contribution is Hydraformer, a modality-aware mixture-of-experts (MoE) fusion block, cascaded with MoE modality encoders and an MoE predictor head. The result sits alongside other recent sequence-first affinity predictors such as AQAffinity and structure-based scoring functions like GatorAffinity, but distinguishes itself by being applied to external screening benchmarks from a single fixed checkpoint.

#Key Features

  • No conformation preparation: Inputs are a receptor sequence and a ligand chemical representation. There is no docking step, no 3D complex, and no per-target structure pipeline to maintain.
  • Cascaded mixture-of-experts: Sparse expert routing is applied at three stages — the modality encoders, the fusion transformer, and the affinity predictor. Ablations show that MoE in any single module yields marginal gains, while cascading across all three compounds them.
  • Modality-aware fusion: Hydraformer gives the prediction token, receptor, and ligand their own routing experts and normalization layers while sharing one large gated feed-forward expert, which aligns modalities implicitly without a dedicated alignment loss.
  • Interpretable routing: Expert-activation patterns differ systematically across Pfam protein families, giving the sparse parameterization a target-family-aware character that can be inspected rather than assumed.
  • Zero-shot screening evaluation: Virtual-screening results come from the PDBbind-trained checkpoint applied directly, with no fine-tuning, hyperparameter adaptation, or exposure to target-domain labels.

#Technical Details

Ligands are encoded by pretrained models spanning SMILES and SELFIES text (MolFormer, ChemBERTa, PepDoRA, MolAI, SELFormer), molecular graphs (Uni-Mol, GeminiMol), and rendered 2D and PyMOL images (ImageMol, MaskMol), capped at 256 tokens or atoms. Receptors use ESM-2 (650M and 3B), ESM-3, SaProt, and ProSST-2048, with embeddings averaged over sequence and then over chains. MoE modules follow the DeepSeek-V3 formulation — a linear router, shared experts, and twelve single-layer routing experts — with noise routing during training and a cross-entropy load-balancing constraint against a uniform target, scaled between 1e-7 and 0.05.

Training uses the EHIGN split of PDBbind v2016 (11,904 training, 1,000 validation complexes) with AdamW at a 1e-4 learning rate, batch size 196 or 256, and early stopping; all runs are repeated across three seeds on a single consumer GPU. On the CASF-2016 core set (285 complexes) the model reaches an average RMSE of 1.161 and a concordance index of 0.826, outperforming all interaction-free baselines and matching state-of-the-art interaction-based methods. Zero-shot enrichment is measured on DUDE-Z (43 targets) and LIT-PCBA (15 targets, ~1:332 active-to-inactive ratio), where HydrAffinity leads on EF5% against PLANET and Glide SP while showing no consistent advantage at EF1%. Performance degrades on the GEMS clean split and on the charge-extreme DUDE-Z Extrema subset, where all tested representations fall to roughly random AUROC.

#Applications

The intended role is an early-stage pre-filter for ultra-large virtual screening: rank a multimillion-compound library at sequence-level cost, retain a large fraction of true actives in the top 5%, and hand a drastically reduced set to slower docking or interaction-based rescoring. The design is deliberate about that positioning — strong EF5% with unremarkable EF1% suits triage rather than final hit prioritization. The same recipe transfers to Davis, KIBA, and BindingDB with only the MoE predictor retained, making it usable for general drug-target affinity regression where structures are unavailable.

#Impact

HydrAffinity is a bioRxiv preprint from a two-author group and has not yet been peer-reviewed; the CASF-2016 comparison is against literature-reported baselines rather than uniformly re-run head-to-head. Its more durable contribution may be the encoder audit, which supplies concrete guidance on which pretrained representations are worth using for affinity tasks and documents that no single ligand encoder wins across LIT-PCBA targets. Reproduction code is public and the trained weights and assembled PDBbind data are deposited on Zenodo, though the code repository carries no license file, which leaves reuse terms for the software itself undefined.

Citation

A Preparation-Free Mixture-of-Experts Framework for Protein-Ligand Affinity Prediction

Bao, H. & Dong, S. (2026) A Preparation-Free Mixture-of-Experts Framework for Protein-Ligand Affinity Prediction. bioRxiv.

DOI: 10.64898/2026.07.24.740495

Recent citations

Papers that recently cited this model.

Not enough citation data yet.

Top citations

The most-cited papers that cite this model.

Not enough citation data yet.

Where to run HydrAffinity

Providers that host HydrAffinity for inference, fine-tuning, or weight download.

No providers recorded yet. Browse all providers

Related models

Models with similar goals, methods, or subject matter.

  • AQAffinity

    SandboxAQ

    Structure-free protein-ligand binding affinity predictor built on OpenFold3 that scores potency from a protein sequence and a ligand SMILES string.

    ProteinSmall molecule
  • GatorAffinity

    University of Florida

    Geometric deep learning scoring function for protein-ligand binding affinity, pretrained on synthetic complexes and fine-tuned on PDBbind structures.

    ProteinSmall molecule
  • HypSeek

    Tsinghua University / University of Electronic Science and Technology of China / Beijing Academy of Artificial Intelligence

    Protein-ligand binding model that embeds ligands, pockets, and sequences in hyperbolic space, unifying virtual screening and affinity ranking.

    ProteinSmall molecule
  • AuroBind

    Shanghai Jiao Tong University / Lingang Laboratory / Sun Yat-sen University / Fudan University / Shanghai AI Laboratory / Shanghai Institute of Materia Medica / MIT / Ningxia Medical University

    Structure-based virtual screening model that jointly predicts protein-ligand complex structures and binding fitness from sequence and SMILES.

    ProteinSmall molecule
  • Zero-Shot Protein-Ligand Binding Site Prediction

    University of Missouri / Politecnico di Milano

    Sequence-based protein-ligand binding site predictor pairing a protein language model with a SMILES chemical language model for zero-shot ligands.

    ProteinSmall molecule
  • SolvCLIP

    The Hong Kong Polytechnic University / Lingnan University / Hong Kong Sanatorium & Hospital

    Protein-ligand interaction model pretrained on solvent-aware conformer ensembles, reaching 97.1% AUC on DUD-E virtual screening.

    Small moleculeProtein
  • BOLD-GPCRs

    Icahn School of Medicine at Mount Sinai

    GPCR ligand bioactivity predictor combining ProteinBERT receptor embeddings with molecular descriptors, spanning the class A receptor family.

    ProteinSmall molecule

Citations

Total Citations0
Influential0
References0

GitHub

Stars0
Forks0
Open Issues0
Contributors1
Last Push6d ago
LanguagePython

Fields of citing research

Not enough data

Openness

bio.rodeo opennessClosed · low usability and reproducibility
38Closed
Usability — can I run it?38
Reproducibility — can I retrain it?44

Tags

binding_affinity_predictionmixture_of_expertsmultimodaltransformervirtual_screening

Resources

GitHub RepositoryResearch PaperDataset