Peptide-protein co-folding and small-molecule docking with an all-atom diffusion transformer. Recovers 59.1% of peptide interfaces to sub-2 Å RMSD.
Co-evolution-driven structure prediction reshaped protein drug discovery, but its accuracy does not carry over to peptide therapeutics. That modality is defined by precisely the features residue-tokenized networks cannot represent: non-canonical amino acids, N-methylation, head-to-tail macrocyclization, disulfide staples, and topologies with no evolutionary record to draw on.
Vilya-2, from Vilya Research, is a diffusion transformer that extends the all-atom representation introduced in Vilya-1 from modeling molecules in isolation to modeling how they bind protein targets. Every input — peptide, small molecule, or protein — is encoded as the same graph of heavy atoms and covalent bonds, with no residue-level tokenization and no molecule-type annotation. That uniformity is what lets the network transfer between molecular classes, so an unnatural macrocycle and a drug-like ligand run through identical machinery.
One trained network therefore spans three tasks usually split across separate tools: peptide-protein co-folding, small-molecule docking, and conformer generation for free molecules. A 2026 preprint positions it as the structure-prediction oracle de novo peptide design pipelines require, and shows it can be fine-tuned to enrich for active compounds in hit-to-lead campaigns.
The architecture revises Vilya-1's by dropping triangle attention in favor of triangle multiplication for pair updates, increasing the ratio of 1D to 2D track updates, and moving dynamic index computation outside the core network so the full graph compiles. Interface prediction is conditioned on a sparse distance matrix over receptor Cα and nucleic acid C4′ coordinates, with inputs cropped to roughly 1,024 atoms centered on the ligand and jittered by a median of 3.2 Å. Training runs in four stages — conformer generation, interface prediction under the same objective, confidence estimation from 100-pose ensembles, and adapter-based activity prediction — on PDB entries released on or before 2021-09-30 plus the CPSea cyclic peptide-protein database.
On Riptides, a curated benchmark of 88 protein-peptide complexes (51 with non-canonical cyclic peptides, all released after 2023-06-01), Vilya-2 recovers 54.1% of interfaces to sub-2 Å backbone RMSD from 100 samples, 59.1% at 1,000 samples, and 38.0% to sub-1 Å. Boltz-2 reaches 40.9% and 22.7% at those thresholds even when given the receptor crystal structure as a template; Schrödinger's MacroDock/Glide reaches 22.0%. For small molecules, success below 2 Å heavy-atom RMSD is 85.2% on PoseBusters, 79.5% on Runs N' Poses, and 76.0% and 71.1% on PoseX self- and cross-docking, ahead of Pearl, Chai-1, and the Boltz series. A retrospective hit-to-lead evaluation reports 3-fold enrichment for actives in the top 20% of the library, against 1.9-fold for ChemProp and 1.3-fold for a fingerprint nearest-neighbor baseline, and near-saturating performance on a quarter of the campaign data.
Vilya-2 is built for the peptide therapeutic workflow: triaging designed macrocycles and stapled miniproteins by predicted bound pose, generating conformational ensembles for binding hypotheses, and scoring compounds during hit-to-lead optimization. Because it docks small molecules competitively, it can serve as one structural engine across both modalities — useful precisely where general-purpose co-folding networks degrade.
Vilya-2's argument is that molecule-type-specific tooling is unnecessary: one chemistry-native representation covers both competitive small-molecule docking and peptide chemistries beyond the reach of co-evolutionary methods. The Riptides release — evaluation code under MIT, PDB-derived data under CC BY 4.0 — gives the field a shared, chemically diverse yardstick for a modality that lacked one. The work is a preprint awaiting peer review, and its benchmarks are retrospective and in silico. The model itself is proprietary: no weights or inference code are released, access runs through the company, and no model card or data card is published.
Papers that recently cited this model.
The most-cited papers that cite this model.
Providers that host Vilya-2 for inference, fine-tuning, or weight download.
No providers recorded yet. Browse all providers
Not enough data