bio.rodeo
ModelsOrganizationsLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
Protein

Efficient Evolution of Human Antibodies from Protein Language Models

Stanford University

Zero-shot antibody affinity maturation using ESM pseudolikelihood scoring. Improves binding up to 160-fold with no antigen-specific training data.

Released: April 2023

This method, developed by Brian Hie, Peter Kim, and colleagues at Stanford University and published in Nature Biotechnology in April 2023, demonstrates that general-purpose protein language models can guide efficient antibody affinity maturation without any antigen-specific training data. Rather than training a task-specific model on binding measurements — which requires a substantial initial dataset and significant experimental investment — the approach repurposes the evolutionary knowledge already encoded in large sequence-based language models to identify mutations that are biologically plausible and likely to improve function.

The core insight is that protein language models, trained on tens of millions of natural protein sequences, implicitly learn what substitutions are evolutionarily tolerated at each position. By scoring candidate mutations using the pseudolikelihood of each amino acid given its sequence context, the method identifies substitutions that nature has already vetted across evolutionary time. This strategy sidesteps the cold-start problem inherent in supervised machine learning-guided directed evolution: it requires nothing more than the wild-type sequence to generate mutation recommendations.

Experimental validation across seven therapeutically relevant antibodies — including those targeting influenza hemagglutinin, Ebola glycoprotein, and SARS-CoV-2 receptor-binding domain — demonstrated that the approach can substantially improve binding affinity while simultaneously maintaining thermostability and, in most cases, improving viral neutralization potency. Critically, the entire laboratory evolution campaign required screening no more than 20 variants per antibody across just two rounds of testing, making the workflow practical for resource-constrained discovery programs.

#Key Features

  • Zero-shot mutation guidance: Recommends beneficial substitutions from the wild-type sequence alone, with no need for initial binding measurements, high-throughput screening data, or task-specific fine-tuning.
  • Consensus ensemble scoring: Aggregates pseudolikelihood scores across multiple protein language model variants (ESM-1b and an ensemble of five ESM-1v models) and accepts only substitutions that exceed a likelihood threshold in at least k models, reducing noise and improving hit rates.
  • Experimentally efficient: Achieves significant affinity gains by screening 20 or fewer variants per antibody in two laboratory rounds — orders of magnitude fewer experiments than conventional directed evolution or deep mutational scanning workflows.
  • Broad applicability across maturation states: Improves both highly mature clinical antibodies (up to sevenfold) and germline-proximal unmatured antibodies (up to 160-fold), demonstrating utility at different stages of antibody development.
  • Multi-property optimization: Improvements in binding affinity correlate strongly with neutralization potency (Spearman r = 0.82), and 21 of 31 tested variants maintained melting temperatures above 70 degrees Celsius, indicating that the evolutionary plausibility filter simultaneously guards against destabilizing mutations.
  • Rapid computation: The scoring pipeline requires less than one second per antibody on GPU hardware, enabling screening of thousands of therapeutic candidates in minutes.

#Technical Details

The method uses ESM-1b (650 million parameters, trained on UniRef50 with approximately 27 million sequences) and an ensemble of five ESM-1v models (each 650 million parameters, trained on UniRef90 with approximately 98 million sequences) as the scoring backbone. For a given wild-type sequence, the pipeline computes the log-likelihood ratio of each possible single-site substitution relative to the wild-type residue, using masked language modeling in the style of BERT. A substitution is recommended only if it exceeds a likelihood ratio threshold alpha and is among the top-scoring changes at that position in at least k of the language models, where k is a tunable stringency parameter.

This consensus strategy outperformed 47 alternative variant-effect predictors on a standardized benchmark and consistently exceeded antibody-specific models such as AbLang and Sapiens, which despite being trained exclusively on antibody sequences failed to match the evolutionary signal captured by general protein models. The authors attribute this counterintuitive result to the much larger and more diverse training corpora of general models, which encode richer epistatic information. The method is also agnostic to antigen identity, antibody class, or target indication, and has been validated across eight diverse protein families beyond antibodies, including beta-lactamase and influenza hemagglutinin.

#Applications

The primary application is antibody affinity maturation during therapeutic development, where the method can be deployed after initial hit identification to improve binding potency without a large experimental dataset. It is particularly valuable for targets where antigen-specific training data are scarce, such as emerging pathogens or novel disease targets. The pipeline is also applicable to the optimization of unmatured germline antibodies, which may retain desirable breadth properties but require affinity improvement for clinical utility. More broadly, the consensus pseudolikelihood scoring framework can be applied to any protein engineering campaign where the goal is to improve an existing function without dramatically altering evolutionary character — including enzyme optimization, cytokine engineering, and the improvement of biosimilar candidates.

#Impact

The study was influential in establishing that general protein language models, without antibody-specific fine-tuning, provide competitive or superior guidance for antibody engineering compared to domain-specific models. It has contributed to a broader shift in the field toward zero-shot and few-shot approaches to protein design, reducing the reliance on large labeled datasets and making machine-learning-guided engineering accessible to smaller laboratories. The paper's emphasis on experimental efficiency — achieving meaningful improvements through minimal screening — directly addresses a practical bottleneck in therapeutic antibody development. Key limitations include the restriction to function-improving rather than function-switching mutations, reduced effectiveness when wild-type sequences already occupy fitness peaks, and the known challenges of generalizing to mutations far outside natural sequence distributions. The experimental code and data are openly available under an MIT license, facilitating adoption across academic and industrial settings.

Citation

Efficient evolution of human antibodies from general protein language models

Hie, B. L., et al. (2023) Efficient evolution of human antibodies from general protein language models. Nature Biotechnology.

DOI: 10.1038/s41587-023-01763-2

Recent citations

Papers that recently cited this model.

  • Engineering Functional CLA-Targeting CAR Approaches for Pancreatic Ductal Adenocarcinoma

    Camille Dourlens, Kim Vanderliek, Lilli Geiger, et al.

    bioRxiv · Jul 2026

    0Influential
  • Breaking the Synthesis Barrier for AI-Designed DNA Libraries

    Scott Sussex, Ema Borevković, F. Lohmann, et al.

    bioRxiv · Jul 2026

    0
  • Biophysical fitness landscape design traps viral evolution.

    Vaibhav Mohanty, Eugene I. Shakhnovich

    Proceedings of the National Academy of Sciences of the United States of America · Jun 2026

    0

Top citations

The most-cited papers that cite this model.

  • ProGen2: Exploring the Boundaries of Protein Language Models

    Erik Nijkamp, Jeffrey A. Ruffolo, Eli N. Weinstein, et al.

    Cell Systems · Jun 2022

    519
  • Bilingual language model for protein sequence and structure

    M. Heinzinger, Konstantin Weissenow, Joaquin Gomez Sanchez, et al.

    bioRxiv · Mar 2024

    317
  • De novo protein design – from new structures to programmable functions

    Tanja Kortemme

    Cell · Feb 2024

    267
  • Machine learning for functional protein design

    Pascal Notin, Nathan J. Rollins, Yarin Gal, et al.

    Nature Biotechnology · Feb 2024

    250
  • Sequence modeling and design from molecular to genome scale with Evo

    Eric Nguyen, Michael Poli, Matthew G. Durrant, et al.

    Science · Nov 2024

    219

Related models

Models with similar goals, methods, or subject matter.

  • CDR-Masked Paired Antibody Language Model

    Boston University

    Paired heavy/light antibody language model fine-tuning ESM-2 and ESM-C with CDR-preferential masking for zero-shot binding affinity embeddings.

    Protein
  • IgLM

    GrayLab

    Generative language model trained on 558 million antibody sequences for infilling-based design of CDR loops and full-length immunoglobulin sequences.

    Protein
  • peleke-1

    Silico Biosciences / Tuple

    Suite of large language models fine-tuned with LoRA to generate antigen-targeted antibody Fv sequences from an antigen and its epitope.

    Protein
  • ReprogBERT

    IBM

    Antibody CDR design model that reprograms a frozen English BERT for sequence infilling, avoiding training a dedicated protein language model.

    Protein
  • SpeciefAI

    University of Edinburgh

    Transformer that generates multi-species antibody and nanobody framework regions at the mRNA level, conditioned on input CDRs, across six species.

    ProteinRNA
  • Walk-Jump Sampling

    Prescient Design / Genentech

    Discrete generative model for antibody protein sequences combining MCMC walks on a smoothed energy landscape with one-step denoising jumps.

    Protein
  • Patch-Centric Conformational B-Cell Epitope Predictor

    Albert Einstein College of Medicine

    Structure-based conformational B-cell epitope predictor that scores local antigen surface patches with ESM-2 embeddings and an ensemble MLP.

    Protein

Citations

Total Citations410
Influential27
References77

GitHub

Stars226
Forks62
Open Issues7
Contributors2
Last Push2y ago
LanguagePython
LicenseMIT

Fields of citing research

  • Computer Science77%
  • Biology70%
  • Medicine68%
  • Chemistry14%
  • Engineering10%
  • Materials Science3%
  • Environmental Science3%
  • Physics1%

Share of papers citing this model.

Openness

bio.rodeo opennessOpen weights · open weights, closed recipe
42Partial
Usability — can I run it?66
Reproducibility — can I retrain it?22
Model Openness Framework
Unclassified
Restrictive license on core components

Tags

antibodydirected_evolutionfoundation_modelsequence_design

Resources

GitHub RepositoryResearch PaperDataset