bio.rodeo
ModelsOrganizationsProvidersLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsProvidersLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
ProteinSmall molecule

Vilya-2

Vilya

Peptide-protein co-folding and small-molecule docking with an all-atom diffusion transformer. Recovers 59.1% of peptide interfaces to sub-2 Å RMSD.

Released: July 2026

Co-evolution-driven structure prediction reshaped protein drug discovery, but its accuracy does not carry over to peptide therapeutics. That modality is defined by precisely the features residue-tokenized networks cannot represent: non-canonical amino acids, N-methylation, head-to-tail macrocyclization, disulfide staples, and topologies with no evolutionary record to draw on.

Vilya-2, from Vilya Research, is a diffusion transformer that extends the all-atom representation introduced in Vilya-1 from modeling molecules in isolation to modeling how they bind protein targets. Every input — peptide, small molecule, or protein — is encoded as the same graph of heavy atoms and covalent bonds, with no residue-level tokenization and no molecule-type annotation. That uniformity is what lets the network transfer between molecular classes, so an unnatural macrocycle and a drug-like ligand run through identical machinery.

One trained network therefore spans three tasks usually split across separate tools: peptide-protein co-folding, small-molecule docking, and conformer generation for free molecules. A 2026 preprint positions it as the structure-prediction oracle de novo peptide design pipelines require, and shows it can be fine-tuned to enrich for active compounds in hit-to-lead campaigns.

#Key Features

  • Chemistry-first all-atom graph: Nodes are heavy atoms and edges covalent bonds, annotated with atomic number, hybridization, formal charge, bond order, and aromaticity — arbitrary synthetic chemistry rather than a 20-residue vocabulary.
  • Ensembles ranked by calibrated confidence: Diverse poses are sampled by diffusion and scored with a predicted per-atom lDDT averaged over ligand atoms; poses above 0.8 on that metric exceed a 60% chance of landing within 2 Å.
  • Template-free interface modeling: Bound poses for cyclic, stapled, and non-canonical peptides are predicted without templates or co-evolutionary input.
  • Docking accuracy that survives novelty: In the lowest training-similarity bin of Runs N' Poses, accuracy holds above 60% where Boltz-2 drops to 20%, and moving from self- to cross-docking costs only about five points.
  • Extrapolation past the training size limit: Trained only on molecules under 128 heavy atoms, it models disulfide-stapled miniproteins two to three times larger, and large Cambridge Structural Database organics to under 1 Å all-atom RMSD.
  • Adapter fine-tuning for potency: Low-rank adapters with gated attention pooling over conformer ensembles turn the pretrained network into an activity ranker.

#Technical Details

The architecture revises Vilya-1's by dropping triangle attention in favor of triangle multiplication for pair updates, increasing the ratio of 1D to 2D track updates, and moving dynamic index computation outside the core network so the full graph compiles. Interface prediction is conditioned on a sparse distance matrix over receptor Cα and nucleic acid C4′ coordinates, with inputs cropped to roughly 1,024 atoms centered on the ligand and jittered by a median of 3.2 Å. Training runs in four stages — conformer generation, interface prediction under the same objective, confidence estimation from 100-pose ensembles, and adapter-based activity prediction — on PDB entries released on or before 2021-09-30 plus the CPSea cyclic peptide-protein database.

On Riptides, a curated benchmark of 88 protein-peptide complexes (51 with non-canonical cyclic peptides, all released after 2023-06-01), Vilya-2 recovers 54.1% of interfaces to sub-2 Å backbone RMSD from 100 samples, 59.1% at 1,000 samples, and 38.0% to sub-1 Å. Boltz-2 reaches 40.9% and 22.7% at those thresholds even when given the receptor crystal structure as a template; Schrödinger's MacroDock/Glide reaches 22.0%. For small molecules, success below 2 Å heavy-atom RMSD is 85.2% on PoseBusters, 79.5% on Runs N' Poses, and 76.0% and 71.1% on PoseX self- and cross-docking, ahead of Pearl, Chai-1, and the Boltz series. A retrospective hit-to-lead evaluation reports 3-fold enrichment for actives in the top 20% of the library, against 1.9-fold for ChemProp and 1.3-fold for a fingerprint nearest-neighbor baseline, and near-saturating performance on a quarter of the campaign data.

#Applications

Vilya-2 is built for the peptide therapeutic workflow: triaging designed macrocycles and stapled miniproteins by predicted bound pose, generating conformational ensembles for binding hypotheses, and scoring compounds during hit-to-lead optimization. Because it docks small molecules competitively, it can serve as one structural engine across both modalities — useful precisely where general-purpose co-folding networks degrade.

#Impact

Vilya-2's argument is that molecule-type-specific tooling is unnecessary: one chemistry-native representation covers both competitive small-molecule docking and peptide chemistries beyond the reach of co-evolutionary methods. The Riptides release — evaluation code under MIT, PDB-derived data under CC BY 4.0 — gives the field a shared, chemically diverse yardstick for a modality that lacked one. The work is a preprint awaiting peer review, and its benchmarks are retrospective and in silico. The model itself is proprietary: no weights or inference code are released, access runs through the company, and no model card or data card is published.

Citation

Preprint

DOI: 10.48550/arXiv.2607.25156

Recent citations

Papers that recently cited this model.

Not enough citation data yet.

Top citations

The most-cited papers that cite this model.

Not enough citation data yet.

Where to run Vilya-2

Providers that host Vilya-2 for inference, fine-tuning, or weight download.

No providers recorded yet. Browse all providers

Related models

Models with similar goals, methods, or subject matter.

  • Vilya-1

    Vilya

    All-atom foundation model for macrocyclic peptide structure prediction, permeability estimation, and de novo design across non-canonical chemistries.

    ProteinSmall molecule
  • Peptide2Mol

    Tsinghua University

    Equivariant diffusion model that converts peptide binders into drug-like small molecules, generating peptidomimetics inside the target protein pocket.

    Small moleculeProtein
  • Pep2Mol

    University of Florida

    Diffusion model for 3D small-molecule design against protein-protein interaction sites, guided by the natural binding peptide or protein partner.

    Small moleculeProtein
  • Pearl

    Genesis Molecular AI

    Protein-ligand cofolding model that predicts 3D complex structures with SO(3)-equivariant diffusion, trained on physics-based synthetic data.

    Protein
  • PeptideCLM-2

    University of Texas at Austin / Novo Nordisk

    Chemical language models pretrained on SMILES for therapeutic peptides, natively representing non-canonical residues, cyclization, and conjugation.

    Small moleculeProtein

Fields of citing research

Not enough data

Openness

bio.rodeo opennessClosed · low usability and reproducibility
20Closed
Usability — can I run it?7
Reproducibility — can I retrain it?34

Tags

diffusiondockingfoundation_modelmacrocyclic_peptidesstructure_prediction

Resources

Research PaperOfficial WebsiteDataset