bio.rodeo
ModelsOrganizationsLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
Protein

ProtGPT2

University of Bayreuth

Autoregressive protein language model based on GPT-2 that generates de novo protein sequences sampling unexplored regions of protein space.

Released: July 2022
Parameters: 738 Million

ProtGPT2 is a generative protein language model developed by Noelia Ferruz, Stefan Schmidt, and Birte Höcker at the University of Bayreuth and published in Nature Communications in July 2022. It applies the autoregressive language modeling paradigm — previously demonstrated on natural language — directly to protein sequences, enabling the sampling of novel proteins that have never existed in nature.

The model addresses a fundamental challenge in protein design: how to efficiently explore the vast space of possible protein sequences without relying on expensive directed evolution campaigns or exhaustive mutagenesis screens. Traditional computational protein design requires expert specification of structural targets and often produces sequences closely related to known proteins. ProtGPT2 takes a different approach: by learning the statistical regularities of natural protein sequences at scale, the model can generate sequences that capture the compositional and structural logic of real proteins while diverging substantially from anything in current databases.

The release of ProtGPT2 coincided with a wave of protein language model research and demonstrated that decoder-only, autoregressive architectures — the same family powering large language models for text — are well-suited to the protein design task without requiring structural supervision.

#Key Features

  • Autoregressive sequence generation: Generates complete protein sequences token-by-token using a causal language modeling objective, enabling open-ended sampling of novel sequences with no structural template required.
  • Large-scale training on UniRef50: Trained on approximately 44.9 million sequences from UniRef50 (2021_04 release), giving the model broad coverage of natural protein diversity across all kingdoms of life.
  • BPE tokenization of amino acid oligomers: Uses a byte-pair encoding (BPE) tokenizer with a vocabulary of 50,256 tokens, where each token corresponds to a frequently reused amino acid oligomer averaging four residues — an approach borrowed from NLP that captures sub-sequence patterns efficiently.
  • High globularity of generated sequences: Disorder predictions show that 88% of ProtGPT2-generated sequences are predicted to be globular, matching the proportion observed in natural proteins and indicating that the model has internalized signals associated with ordered structure.
  • Exploration of novel protein space: Sequence similarity searches confirm that ProtGPT2 outputs are distantly related to natural proteins, and similarity network analyses demonstrate that the model samples regions of sequence space not occupied by known proteins.
  • AlphaFold-compatible output: Structure prediction of ProtGPT2 sequences using AlphaFold 2 yields well-folded, non-idealized structures featuring topologies not found in current structure databases, providing computational evidence that the generated sequences encode viable protein folds.

#Technical Details

ProtGPT2 is a 738-million parameter decoder-only transformer based on the GPT-2 architecture, scaled to 36 layers with a model dimensionality of 1,280. The architecture is identical in design to GPT-2 XL but adapted for protein sequence input via the BPE tokenizer trained on oligomer vocabularies from protein space. Training used a causal language modeling objective — predicting the next token given all preceding tokens — which makes the model naturally suited to sequential generation.

The model was trained on UniRef50 (April 2021 release), a non-redundant protein sequence database clustering sequences at 50% identity. Approximately 44.9 million sequences were used for training, with 4.9 million held out for evaluation. Sequences are delimited by special tokens marking the beginning and end of each protein, allowing the model to generate complete sequences of variable length. No structural or functional labels were used during training; all information is derived from the sequence distribution alone. The model was made publicly available through HuggingFace, where sequences can be generated in seconds on consumer hardware.

#Applications

ProtGPT2 is primarily used as a starting point for de novo protein design workflows. Researchers can sample thousands of candidate sequences rapidly and then filter by predicted structural quality (using AlphaFold 2 pLDDT scores), predicted function, or experimental fitness. The model is also useful for fine-tuning on specific protein families — its pretrained representations can be adapted to generate sequences constrained to a particular fold or functional class with relatively small domain-specific datasets. Beyond generation, ProtGPT2 has been used as a perplexity-based scoring function to evaluate the naturalness of engineered sequences, analogous to masked language model scoring in models like ESM. Wet-lab validation in the original study confirmed that ProtGPT2-generated sequences can be expressed in experimental systems, establishing practical relevance beyond computational benchmarks.

#Impact

ProtGPT2 was one of the first demonstrations that large autoregressive language models, trained without any structural supervision, could generate protein sequences with the hallmarks of natural proteins at scale. Published alongside contemporaneous work such as ProGen and ESM, it helped establish the generative protein language model as a distinct and productive research direction. The paper has attracted substantial citations and the HuggingFace model has seen broad community adoption for sequence generation and fine-tuning tasks. A notable limitation is that ProtGPT2, like all sequence-only generative models, lacks explicit structural control — there is no mechanism to steer generation toward a specific topology or binding site without additional downstream filtering or fine-tuning. Subsequent models such as EvoDiff, RFdiffusion, and Chroma have addressed structural controllability, but ProtGPT2 remains a widely used baseline for unconditional protein sequence generation due to its accessibility and strong sequence-level properties.

Citation

ProtGPT2 is a deep unsupervised language model for protein design

Ferruz, N., et al. (2022) ProtGPT2 is a deep unsupervised language model for protein design. Nature Communications.

DOI: 10.1038/s41467-022-32007-7

Recent citations

Papers that recently cited this model.

  • ProtAug: An Empirical Investigation of pLM-Guided Data Augmentation for Protein Sequence Prediction Tasks

    Zhuoyang Chen, Ruoqing Wang, Qiong Luo

    bioRxiv · Jul 2026

    0
  • Applications and limitations of AI tools in enzyme design

    Rosa Teijeiro-Juiz, Nina Egeler, Grzegorz Jamróg, et al.

    Protein Science · Jul 2026

    1
  • PCTPS: Accurate prediction of phase-separated proteins using protein language model embeddings and a CTN-KAN framework.

    Taigang Liu, Yuxin Xia, Xuqi Sun, et al.

    Journal of Molecular Graphics and Modelling · Jul 2026

    0

Top citations

The most-cited papers that cite this model.

  • ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope

    P. Ray

    Internet of Things and Cyber-Physical Systems · Apr 2023

    2.1K
  • De novo design of protein structure and function with RFdiffusion

    Joseph L. Watson, David Juergens, N. Bennett, et al.

    Nature · Jul 2023

    1.1KInfluential
  • Large language models generate functional protein sequences across diverse families

    Ali Madani, Ben Krause, E. Greene, et al.

    Nature Biotechnology · Jan 2023

    1K
  • ProGen2: Exploring the Boundaries of Protein Language Models

    Erik Nijkamp, Jeffrey A. Ruffolo, Eli N. Weinstein, et al.

    Cell Systems · Jun 2022

    519
  • HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution

    Eric Nguyen, Michael Poli, Marjan Faizi, et al.

    Neural Information Processing Systems · Jun 2023

    487

Related models

Models with similar goals, methods, or subject matter.

  • ProGen2

    Salesforce

    Protein language models from 151M to 6.4B parameters, trained on over a billion sequences for sequence generation and zero-shot fitness prediction.

    Protein
  • ProGen3

    Profluent

    Sparse mixture-of-experts autoregressive protein language model family pretrained on 1.5 trillion amino acid tokens with compute-optimal scaling.

    Protein
  • sm_protgpt2

    University of Naples Federico II / University of Bern

    Three fixed ProtGPT2 fine-tunes specialized for metalloprotein generation, trained on ProteinMPNN-derived synthetic sequences.

    Protein
  • Prot2Token

    University of Missouri

    Multi-task protein framework recasting function, binding site, and structure prediction as autoregressive next-token prediction over ESM2 embeddings.

    Protein
  • ProtTrans

    Rostlab

    Suite of six protein language models, including ProtBERT and ProtT5, trained on up to 393 billion amino acids without multiple sequence alignments.

    Protein
  • Dayhoff Atlas

    Microsoft Research

    Protein language models trained on billions of natural and synthetic sequences for de novo design and zero-shot mutation-effect prediction.

    Protein
  • Pinal

    Westlake University

    De novo protein design from natural language: a 16B-parameter framework turning text descriptions into sequences via structure-conditioned generation.

    Protein
  • CodeFP

    PharMolix Inc. / Tsinghua University

    Co-generative protein language model decoding sequence and structure tokens together from GO functional annotations for de novo protein design.

    Protein

Citations

Total Citations869
Influential45
References82

HuggingFace

Downloads8.2K
Likes116
Last Modified1mo ago
Pipelinetext-generation

Fields of citing research

  • Computer Science40%
  • Biology33%
  • Medicine27%
  • Chemistry7%
  • Engineering4%
  • Materials Science2%
  • Environmental Science1%
  • Physics1%

Share of papers citing this model.

Openness

bio.rodeo opennessOpen weights · open weights, closed recipe
54Partial
Usability — can I run it?69
Reproducibility — can I retrain it?23
Model Openness Framework
Class III
Open Model

Tags

foundation_modelgenerativeprotein_design

Resources

Research PaperHuggingFace ModelDataset