bio.rodeo
ModelsOrganizationsLeaderboardAbout
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
Protein

Dayhoff Atlas

Microsoft Research

Protein language models trained on billions of natural and synthetic sequences for de novo design and zero-shot mutation-effect prediction.

Released: July 2025
Parameters: 3 Billion

The Dayhoff Atlas is a family of generative protein language models from Microsoft Research, released alongside two large protein-sequence corpora that were assembled to train them. Its central question is whether the diversity of the training data, rather than model size alone, is the limiting factor in protein generation. To test this, the team built the largest open dataset of natural protein sequences reported to date and paired it with a synthetic dataset derived from de novo designed structures, then trained models that learn from both individual proteins and sets of evolutionarily related homologs at scale.

Named for Margaret Dayhoff, a founder of computational protein science, the Atlas addresses a practical gap in the protein language model landscape. Models such as ESM and ProtGPT2 are trained largely on clustered natural sequences and treat proteins one at a time. Dayhoff instead combines single sequences with unrolled multiple sequence alignments in one model, letting it condition generation on evolutionary context. The result is a set of frozen checkpoints that perform mutation-effect scoring, unconditional and conditioned sequence generation, and structural motif scaffolding without task-specific fine-tuning.

The Atlas is released as a bioRxiv preprint with open weights and code, making both the models and their training corpora available to the community.

#Key Features

  • Hybrid state-space architecture: Dayhoff combines Mamba state-space layers, transformer self-attention, and mixture-of-experts modules, giving efficient handling of the very long contexts needed to process concatenated homolog sets while retaining explicit sequence lookup.
  • Evolutionary context in a single model: It is the first protein language model to jointly train on individual sequences and sets of evolutionarily related homologs at scale, enabling homolog-conditioned generation from a single set of weights.
  • Zero-shot mutation-effect prediction: The models score the fitness impact of point mutations directly from pretrained likelihoods, evaluated on ProteinGym without supervised fine-tuning.
  • Conditioned generation and scaffolding: Checkpoints support unconditional design, guided generation within a specified protein family, and scaffolding of structural motifs by conditioning on evolutionary or structural context.
  • Order-agnostic generation: The models generate in both N-to-C and C-to-N directions and perform order-agnostic infilling, supporting flexible design workflows.

#Technical Details

Dayhoff is released in 170-million-parameter (six variants) and 3-billion-parameter (three variants) sizes, with the flagship Dayhoff-3b-GR-HM-c trained across natural sequences, homolog sets, and synthetic backbone-derived sequences. Training draws on GigaRef, a corpus of 3.34 billion protein sequences across 1.70 billion clusters; BackboneRef, 46 million synthetic sequences predicted from 240,811 de novo designed backbones; and OpenProteinSet, roughly 16 million multiple sequence alignments. The models are evaluated on ProteinGym for mutation-effect prediction and on MotifBench and RFDiffusion-style motif-scaffolding tasks for structure-conditioned generation. Code and weights are distributed under the MIT license through GitHub, HuggingFace, and Azure AI Foundry.

#Applications

The Dayhoff models serve protein engineers and computational biologists who need to generate candidate sequences, rank mutations, or design proteins around a fixed functional motif. Because scoring and generation run from frozen checkpoints, the models slot into design pipelines without retraining: a group can screen mutational libraries for likely fitness effects, sample novel members of an enzyme family, or scaffold a binding motif into new sequence contexts. The accompanying GigaRef and BackboneRef datasets are independently useful for training and benchmarking other protein models.

#Impact

By decoupling data diversity from parameter count and releasing the underlying corpora openly, the Dayhoff Atlas gives the field a controlled way to study how training-set breadth shapes generative protein models. Its hybrid state-space design demonstrates a route to processing long evolutionary contexts more efficiently than attention alone, and its joint treatment of single sequences and homolog sets broadens what a single protein language model can condition on. As a preprint awaiting peer review, its benchmark comparisons are in-silico; experimental validation of designed sequences remains future work, but the open release of models, code, and data lowers the barrier for others to build on the approach.

Citation

The Dayhoff Atlas: scaling sequence diversity for improved protein generation

Preprint

Yang, K. K., et al. (2026) The Dayhoff Atlas: scaling sequence diversity for improved protein generation. bioRxiv.

DOI: 10.1101/2025.07.21.665991

Recent citations

Papers that recently cited this model.

  • ProtGPT3: an Open-source family of Promptable and Aligned Protein Language Models

    Michele Garibbo, Gerard Boxó, Filippo Stocco, et al.

    bioRxiv · Jun 2026

    0Influential
  • FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications

    Kieran Didi, Sarah Alamdari, Alex X. Lu, et al.

    bioRxiv · May 2026

    0
  • ProteinJEPA: Latent prediction complements protein language models

    Dan Ofer, Dafna Shahaf, M. Linial

    May 2026

    0

Top citations

The most-cited papers that cite this model.

  • Illuminating the universe of enzyme catalysis in the era of artificial intelligence.

    Jason Yang, Francesca-Zhoufan Li, Yueming Long, et al.

    Cell Systems · Aug 2025

    9
  • Steering generative models for protein design: Aligning and conditioning strategies.

    Filippo Stocco, Michele Garibbo, Noelia Ferruz

    Current Opinion in Structural Biology · Nov 2025

    5
  • Trainable subnetworks reveal insights into structure knowledge organization in protein language models

    R. Vinod, A. Amini, L. Crawford, et al.

    bioRxiv · Dec 2025

    2
  • “Visualize, Explore, and Select”: A Protein Language Model-based Approach Enabling Navigation of Protein Sequence Space for Enzyme Discovery and Mining

    Felix Moorhoff, David Medina-Ortiz, Alicja Kotnis, et al.

    bioRxiv · Mar 2026

    1
  • ProtGPT3: an Open-source family of Promptable and Aligned Protein Language Models

    Michele Garibbo, Gerard Boxó, Filippo Stocco, et al.

    bioRxiv · Jun 2026

    0Influential

Related models

Models with similar goals, methods, or subject matter.

  • ProtGPT2

    University of Bayreuth

    Autoregressive protein language model based on GPT-2 that generates de novo protein sequences sampling unexplored regions of protein space.

    Protein
  • ProGen3

    Profluent

    Sparse mixture-of-experts autoregressive protein language model family pretrained on 1.5 trillion amino acid tokens with compute-optimal scaling.

    Protein
  • ProGen2

    Salesforce

    Protein language models from 151M to 6.4B parameters, trained on over a billion sequences for sequence generation and zero-shot fitness prediction.

    Protein
  • AIDO.Protein

    genbio.ai

    Mixture-of-experts protein language model scaling to 16 billion parameters, applied to variant effect prediction and de novo protein design.

    Protein
  • ProFam

    University College London / Technical University of Munich

    Protein-family language model trained on unaligned homolog sets for zero-shot variant fitness prediction and design. ProFam-1 holds 251M parameters.

    Protein

Citations

Total Citations11
Influential1
References0

GitHub

Stars99
Forks6
Open Issues0
Contributors2
Last Push11d ago
LanguagePython
LicenseMIT

Fields of citing research

  • Computer Science100%
  • Biology91%
  • Medicine45%
  • Mathematics9%
  • Chemistry9%

Share of papers citing this model.

Openness

bio.rodeo opennessFully open · usable and reproducible
96Open
Usability — can I run it?100
Reproducibility — can I retrain it?92
Model Openness Framework
Class II
Open Tooling

Tags

language_modelmixture_of_expertsprotein_designstate_space_modelvariant_effect_prediction

Resources

GitHub RepositoryResearch PaperOfficial WebsiteHuggingFace ModelDataset