All-atom foundation model for macrocyclic peptide structure prediction, permeability estimation, and de novo design across non-canonical chemistries.
Macrocyclic peptides occupy a therapeutically attractive middle ground between small molecules and biologics: they are large enough to engage difficult protein-protein interfaces yet compact enough to be synthesized chemically and, in favorable cases, to cross cell membranes. Modeling them computationally is hard because they routinely incorporate non-canonical amino acids, N-methylations, D-stereochemistry, cyclization backbones, and other chemistries that fall outside the training distribution of conventional protein structure predictors. Existing tools tend to be narrow in chemical scope and generalize poorly across the synthetically accessible design space.
Vilya-1 is an all-atom deep learning foundation model that addresses two coupled problems in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries, and predicting developability properties such as membrane permeability. Rather than fitting a bespoke model per target or per chemical series, it operates on a uniform all-atom representation and is trained once, then applied across a broad set of macrocycles composed of canonical and non-canonical residues without per-molecule retraining.
The model was introduced in a 2026 preprint by researchers at Vilya, a macrocyclic peptide drug discovery company spun out of the Institute for Protein Design. It spans both predictive tasks — conformer and ensemble generation, permeability estimation — and generative tasks, producing novel macrocycles with tailored chemical and structural profiles.
Vilya-1 is trained on heterogeneous structural datasets spanning diverse macrocycle topologies and chemical classes, and operates entirely at the all-atom level. In the accompanying preprint, the authors report improved geometric accuracy relative to physics-based conformer generators, co-folding networks, and deep-learning conformer generators, with the developers describing more than a doubling of structure-recapitulation performance over the next best method. The model additionally couples structure prediction with property prediction, allowing permeability to be estimated from generated conformations and fine-tuning toward that objective. The specific network architecture family, parameter count, and full training-set composition are not disclosed in the preprint.
Vilya-1 targets the macrocyclic peptide drug discovery workflow: proposing and triaging design hypotheses, generating conformational ensembles to reason about binding and membrane behavior, prioritizing candidates by predicted permeability, and generating novel macrocycles with desired properties. The company reports using the model across internal therapeutic programs. Its ability to model non-canonical chemistries makes it useful precisely where general-purpose protein structure predictors break down, giving medicinal chemists and computational designers a way to explore synthetically accessible chemical space that lies outside natural amino acid space.
Vilya-1 extends the structure-prediction and design paradigm that reshaped protein science into the non-canonical, chemically diverse territory of macrocycles, a modality of growing pharmaceutical interest. By unifying conformer generation, permeability prediction, and generative design in one all-atom model, it addresses a gap that sequence-centric protein tools leave open. The work is a preprint awaiting peer review, and its in-silico benchmarks have not yet been accompanied by public prospective wet-lab validation. Vilya-1 is a proprietary commercial model: the weights and code are not publicly released, and access is gated through the company rather than distributed openly, with no public model card or data card available.
Sturmfels, V. R. P., et al. (2026) Vilya-1: An all-atom foundation model for macrocycle structure prediction and design.
DOI: 10.48550/arXiv.2607.09998Papers that recently cited this model.
The most-cited papers that cite this model.
Not enough data