Protein language models trained on billions of natural and synthetic sequences for de novo design and zero-shot mutation-effect prediction.
The Dayhoff Atlas is a family of generative protein language models from Microsoft Research, released alongside two large protein-sequence corpora that were assembled to train them. Its central question is whether the diversity of the training data, rather than model size alone, is the limiting factor in protein generation. To test this, the team built the largest open dataset of natural protein sequences reported to date and paired it with a synthetic dataset derived from de novo designed structures, then trained models that learn from both individual proteins and sets of evolutionarily related homologs at scale.
Named for Margaret Dayhoff, a founder of computational protein science, the Atlas addresses a practical gap in the protein language model landscape. Models such as ESM and ProtGPT2 are trained largely on clustered natural sequences and treat proteins one at a time. Dayhoff instead combines single sequences with unrolled multiple sequence alignments in one model, letting it condition generation on evolutionary context. The result is a set of frozen checkpoints that perform mutation-effect scoring, unconditional and conditioned sequence generation, and structural motif scaffolding without task-specific fine-tuning.
The Atlas is released as a bioRxiv preprint with open weights and code, making both the models and their training corpora available to the community.
Dayhoff is released in 170-million-parameter (six variants) and 3-billion-parameter (three variants) sizes, with the flagship Dayhoff-3b-GR-HM-c trained across natural sequences, homolog sets, and synthetic backbone-derived sequences. Training draws on GigaRef, a corpus of 3.34 billion protein sequences across 1.70 billion clusters; BackboneRef, 46 million synthetic sequences predicted from 240,811 de novo designed backbones; and OpenProteinSet, roughly 16 million multiple sequence alignments. The models are evaluated on ProteinGym for mutation-effect prediction and on MotifBench and RFDiffusion-style motif-scaffolding tasks for structure-conditioned generation. Code and weights are distributed under the MIT license through GitHub, HuggingFace, and Azure AI Foundry.
The Dayhoff models serve protein engineers and computational biologists who need to generate candidate sequences, rank mutations, or design proteins around a fixed functional motif. Because scoring and generation run from frozen checkpoints, the models slot into design pipelines without retraining: a group can screen mutational libraries for likely fitness effects, sample novel members of an enzyme family, or scaffold a binding motif into new sequence contexts. The accompanying GigaRef and BackboneRef datasets are independently useful for training and benchmarking other protein models.
By decoupling data diversity from parameter count and releasing the underlying corpora openly, the Dayhoff Atlas gives the field a controlled way to study how training-set breadth shapes generative protein models. Its hybrid state-space design demonstrates a route to processing long evolutionary contexts more efficiently than attention alone, and its joint treatment of single sequences and homolog sets broadens what a single protein language model can condition on. As a preprint awaiting peer review, its benchmark comparisons are in-silico; experimental validation of designed sequences remains future work, but the open release of models, code, and data lowers the barrier for others to build on the approach.
Yang, K. K., et al. (2026) The Dayhoff Atlas: scaling sequence diversity for improved protein generation. bioRxiv.
DOI: 10.1101/2025.07.21.665991Papers that recently cited this model.
Michele Garibbo, Gerard Boxó, Filippo Stocco, et al.
bioRxiv · Jun 2026
Kieran Didi, Sarah Alamdari, Alex X. Lu, et al.
bioRxiv · May 2026
Dan Ofer, Dafna Shahaf, M. Linial
May 2026
The most-cited papers that cite this model.
Jason Yang, Francesca-Zhoufan Li, Yueming Long, et al.
Cell Systems · Aug 2025
Filippo Stocco, Michele Garibbo, Noelia Ferruz
Current Opinion in Structural Biology · Nov 2025
R. Vinod, A. Amini, L. Crawford, et al.
bioRxiv · Dec 2025
Felix Moorhoff, David Medina-Ortiz, Alicja Kotnis, et al.
bioRxiv · Mar 2026
Michele Garibbo, Gerard Boxó, Filippo Stocco, et al.
bioRxiv · Jun 2026
Share of papers citing this model.