bio.rodeo
ModelsOrganizationsLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
Protein

Boltz-1

MIT

Open-source structure prediction model for proteins, nucleic acids, and small molecules, trained on public data to AlphaFold3-level accuracy.

Released: November 2024

Boltz-1 is an open-source deep learning model developed at MIT for predicting the three-dimensional structures of biomolecular complexes, including proteins, nucleic acids, and small molecules. Released in November 2024 by researchers Jeremy Wohlwend, Gabriele Corso, Saro Passaro, and colleagues from MIT's Jameel Clinic under the guidance of professors Regina Barzilay and Tommi Jaakkola, Boltz-1 is notable for being the first fully open model to achieve accuracy on par with AlphaFold3 — the state-of-the-art system from Google DeepMind that had previously been available only under a restrictive non-commercial license.

The model's central contribution to the field is not just technical performance, but democratization. AlphaFold3 represented a significant advance in modeling biomolecular interactions involving proteins, DNA, RNA, and small molecule ligands simultaneously, but its inaccessibility to commercial and independent researchers created a meaningful gap in the ecosystem. Boltz-1 closes that gap by releasing training code, model weights, datasets, and benchmarks under the MIT license, making frontier-level structural biology AI available to the global research community without restriction.

Boltz-1 follows the general architectural framework established by AlphaFold3 but introduces several meaningful innovations in multiple sequence alignment (MSA) pairing, training-time structure cropping, pocket-conditioned prediction, and inference-time steering. The entire training pipeline relies exclusively on publicly available data from the Protein Data Bank, UniProt, and ChEMBL, demonstrating that AlphaFold3-class performance does not require proprietary datasets.

#Key Features

  • AlphaFold3-level accuracy: Achieves performance on par with AlphaFold3 across diverse benchmarks, including protein-ligand docking (LDDT-PLI of 65% on CASP15 targets) and protein-protein complex prediction (DockQ > 0.23 on 83% of CASP15 targets).
  • Fully open-source: Model weights, training code, inference code, datasets, and benchmarks are all released under the MIT license, enabling unrestricted academic and commercial use.
  • Broad molecular scope: Predicts structures of proteins, DNA, RNA, and small molecule ligands as well as multi-chain complexes, supporting the full range of biomolecular interaction types relevant to drug discovery and protein engineering.
  • Pocket-conditioned prediction: Supports user-defined binding pocket constraints at inference time, allowing researchers to guide predictions toward known or hypothesized binding sites and improving relevance for structure-based drug design workflows.
  • Boltz-steering inference technique: A novel inference-time steering approach that corrects hallucinations and physically implausible predictions by applying guidance during the diffusion sampling process.
  • Efficient training design: Introduces improved MSA pairing algorithms and a unified cropping strategy (with crop sizes of 384 and 512 tokens) that substantially reduce computational cost compared to comparable models without sacrificing accuracy.

#Technical Details

Boltz-1 is built around a diffusion-based architecture that closely follows the AlphaFold3 framework, with a trunk network that processes paired sequence and structural representations and a diffusion module that generates atomic coordinates. The model was trained entirely on open data: pre-processed Protein Data Bank (PDB) structures with pre-computed MSAs containing up to 4,096 sequences per chain, along with ligand information sourced from ChEMBL and chemical databases. Key architectural modifications relative to AlphaFold3 include revised MSA pairing algorithms for handling heteromeric complexes, changes to representation flow within the trunk, and a reworked confidence model that frames confidence estimation as a fine-tuning task on the trunk layers rather than a separate head.

On the CASP15 benchmark (66 targets from the 2022 competition), Boltz-1 demonstrates strong performance across modalities: a median LDDT-PLI of 65% for protein-ligand interactions versus 40% for Chai-1, and 83% of protein-protein targets with DockQ > 0.23 versus 76% for Chai-1. RNA prediction performance achieves a median LDDT of 0.54 on CASP15 RNA targets. The Boltz-steering technique further improves output quality at inference time without retraining, by applying constraint-based guidance during diffusion sampling to eliminate non-physical bond geometries and clashes.

#Applications

Boltz-1 is designed to serve the full spectrum of structural biology use cases that previously required access to commercial or institutional tools. Drug discovery teams can use it for structure-based virtual screening, predicting protein-ligand poses for hit identification and lead optimization. The pocket-conditioning feature makes it directly applicable to fragment-based drug design, where partial information about a binding site is used to guide complex structure prediction. Protein engineers can use Boltz-1 to predict heteromeric complex structures, assess the effects of mutations on binding interfaces, or validate designed protein-protein interactions prior to experimental synthesis. Academic researchers benefit from the transparent training pipeline, which supports reproducibility, ablation studies, and further model development in ways that closed systems cannot.

#Impact

Boltz-1 arrived at a moment when the structural biology community was acutely aware of the tension between scientific capability and access. AlphaFold3's non-commercial license had excluded large portions of the research ecosystem from using the most powerful available tool for biomolecular complex prediction. Boltz-1 resolved this by demonstrating that AlphaFold3-level accuracy is achievable with open data and open methods, establishing a new baseline for what the community can expect from open-source tools. The MIT license enables integration into commercial drug discovery platforms, downstream model development, and open-science initiatives. The concurrent release of training data on AWS Open Data further lowers the barrier for groups wishing to replicate or extend the work. A key limitation is that, as a diffusion model, Boltz-1 can still produce hallucinated or non-physical structures for difficult targets — a challenge the Boltz-steering technique partially addresses but does not fully eliminate. As of early 2025, the model has seen rapid adoption in both academic and commercial contexts, with active community development continuing through the public GitHub repository.

Citation

Boltz-1: Democratizing Biomolecular Interaction Modeling

Preprint

Wohlwend, J., Corso, G., Passaro, S., et al. (2024). Boltz-1: Democratizing Biomolecular Interaction Modeling. bioRxiv.

DOI: 10.1101/2024.11.19.624167

Recent citations

Papers that recently cited this model.

  • Characterising AlphaFold 3’s ability to predict T cell antigen specificity

    Benjamin McMaster, Ali El Moselhy, Ilija Ilievski, et al.

    bioRxiv · Jul 2026

    0
  • Capabilities, specificity gaps and training-data dependence of AlphaFold3 across diverse application areas

    O. Follonier, Yan Liu, Pablo Campomanes, et al.

    bioRxiv · Jul 2026

    0
  • Benchmarking AI Protein Structure Predictors Reveals a Persistent Bias in Multi-State Proteins

    Muhui Ye, Yu-Hong Wang, M. Brogi, et al.

    bioRxiv · Jul 2026

    0

Top citations

The most-cited papers that cite this model.

  • Atom-level enzyme active site scaffolding using RFdiffusion2

    Woody Ahern, Jason Yim, D. Tischer, et al.

    bioRxiv · Apr 2025

    89
  • AlphaFold3: An Overview of Applications and Performance Insights

    Marios G. Krokidis, Dimitrios E. Koumadorakis, Konstantinos Lazaros, et al.

    International Journal of Molecular Sciences · Apr 2025

    85
  • Have protein-ligand cofolding methods moved beyond memorisation?

    Peter Škrinjar, Jérôme Eberhardt, G. Tauriello, et al.

    bioRxiv · Aug 2025

    71
  • AI-driven protein design

    Huan Yee Koh, Yi Zheng, Maddie Yang, et al.

    Nature Reviews Bioengineering · Sep 2025

    49
  • Investigating whether deep learning models for co-folding learn the physics of protein-ligand interactions

    Matthew Masters, Amr H. Mahmoud, Markus A. Lill

    Nature Communications · Oct 2025

    46Influential

Related models

Models with similar goals, methods, or subject matter.

  • Boltz-2

    MIT CSAIL / Recursion Pharmaceuticals

    Open model that jointly predicts biomolecular structure and small-molecule binding affinity, approaching FEP+ accuracy in seconds on a single GPU.

    Protein
  • BoltzMol-1

    Boltz

    Small-molecule hit-discovery pipeline using Boltz-2 co-folding and affinity prediction to rank in-stock compounds or make-on-demand chemical space.

    Small moleculeProtein
  • BoltzGen

    MIT

    All-atom generative model for de novo protein and peptide binder design against diverse biomolecular targets, wet-lab validated across 26 targets.

    ProteinSmall molecule
  • BoltzProt-1

    Boltz

    De novo protein binder and nanobody design pipeline that ranks candidates by a protein-protein interaction model rather than structural confidence.

    Protein
  • OpenFold3

    Aqlaboratory / Lawrence Livermore National Laboratory / Seoul National University

    Open-source Apache-2.0 reproduction of AlphaFold3 that predicts all-atom structures of proteins, RNA, DNA, small molecules, and their complexes.

    ProteinRNASmall molecule

Citations

Total Citations431
Influential52
References29

GitHub

Stars4.1K
Forks863
Open Issues159
Contributors63
Last Push1mo ago
LanguagePython
LicenseMIT

HuggingFace

Downloads0
Likes49
Last Modified1y ago

Fields of citing research

  • Biology68%
  • Computer Science68%
  • Medicine49%
  • Chemistry35%
  • Physics6%
  • Environmental Science5%
  • Materials Science4%
  • Engineering3%

Share of papers citing this model.

Openness

bio.rodeo opennessFully open · usable and reproducible
97Open
Usability — can I run it?100
Reproducibility — can I retrain it?95
Model Openness Framework
Class I
Open Science

Tags

drug_discoveryfoundation_modelprotein_protein_interactionstructure_prediction

Resources

GitHub RepositoryResearch PaperHuggingFace ModelDocumentationDataset