bio.rodeo
ModelsOrganizationsProvidersLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsProvidersLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
models / protein / mage-antibody
Protein
Vanderbilt University Medical CenterUniversity of Texas at AustinKarolinska InstitutetCleveland ClinicNational Institute of Allergy and Infectious DiseasesGriffith UniversityUniversity of WashingtonVanderbilt UniversityReleased December 2024

MAGE (Monoclonal Antibody GEnerator)

Protein language model that generates paired heavy and light chain human antibodies from an antigen prompt, with binders validated in vitro.

The short version

  • —Designs antibodies against an emerging pathogen with no template antibody or complex structure
  • —One antigen sequence prompt returns a complete paired VH and VL variable region
  • —Generated binders neutralized SARS-CoV-2, H5N1 influenza, and RSV-A in cell assays
  • —Sequence edits land throughout the framework regions, beyond the CDR loops
67Openness5Citations
0HF downloads
79GitHub stars
Apache-2.0License

Where to run it

No providers recorded yet. Browse all providers

Monoclonal antibody discovery still begins with biology: immunize or sample a donor, sort antigen-specific B cells, screen, and iterate. MAGE (Monoclonal Antibody GEnerator) replaces that first step with generation. Prompted with the amino acid sequence of an antigen, it emits a complete paired antibody variable region — heavy and light chain together — predicted to bind that target. Nothing about the antibody is supplied as input: no seed CDR, no parent clone, no structure of the complex.

That framing separates MAGE from the two dominant computational strategies. Sequence-based protein language models generate antibody-like sequences target-agnostically, leaving unanswered what they bind. Structure-based design conditions on the antigen but depends on solved or predicted antibody-antigen complexes, a thin data regime. MAGE instead learns the association between antigen sequence and binding antibody sequence directly, so one checkpoint serves any antigen without per-target retraining.

MAGE was developed in the Georgiev lab at Vanderbilt University Medical Center with collaborators at UT Austin, Karolinska Institutet, Cleveland Clinic, NIAID, Griffith University, and the University of Washington, and released as a preprint in December 2024. It is a fine-tune of ProGen2-base rather than a model trained from scratch, inheriting general protein sequence knowledge and specializing it for the antibody task.

#Key Features

  • Antigen-prompted paired-chain output: A single forward pass returns both variable heavy and variable light sequences for the prompted antigen, rather than a heavy chain that must later be paired or a CDR loop grafted into an existing scaffold.
  • Template-free design: No starting antibody sequence and no antigen-antibody structure are required, which makes the model usable for targets that have neither.
  • Wet-lab validated binders: Designs were expressed and tested against SARS-CoV-2 RBD, H5 hemagglutinin, and RSV-A prefusion F, with neutralizing antibodies confirmed for all three targets and a cryo-EM structure solved for two RSV designs.
  • Repertoire-scale diversity: RBD prompts produced 322 distinct heavy-light V gene pairings across 969 filtered sequences, with the most common pairing accounting for only 13.9% of the pool.
  • Edits beyond the CDRs: Generated sequences differ from their nearest training antibody throughout the variable region, including framework positions, rather than varying only the CDR3 loop.

#Technical Details

MAGE fine-tunes ProGen2-base, a 764M-parameter decoder-only protein transformer, on a curated database of antibody-antigen sequence pairs assembled from the literature and public databases, augmented with LIBRA-seq data generated for the study — a panel of 18 antigens screened against PBMCs from 20 donors spanning HIV-infected, influenza-vaccinated, COVID-19 convalescent, and healthy groups. Training used four V100 GPUs; generation takes roughly 15 seconds per antibody on an A6000. Outputs are filtered with ANARCI under IMGT numbering and scored for humanness with BioPhi OASis, which retained 969 of 1,000 RBD-prompted sequences.

Experimental validation covered three antigens of decreasing training representation. Of 20 RBD designs tested, 9 bound by ELISA and 8 of those by biolayer interferometry, five with apparent nanomolar to sub-nanomolar affinity; RBD-409 neutralized index SARS-CoV-2 pseudovirus at an IC50 of 6.7 ng/mL and retained potency against Gamma, Delta, and several Omicron variants. For an H5 hemagglutinin from A/Texas/37/2024 — a sequence absent from training, though 472 antibodies against a related H5N1 strain sharing 91.5% identity were present — 5 of 18 designs bound strongly and all five neutralized three influenza strains, two at IC50 below 100 ng/mL. For RSV-A prefusion F, 7 of 23 designs bound and 3 neutralized; a 3.4 Å cryo-EM reconstruction showed Fabs RSV-2245 and RSV-3301 engaging antigenic sites V and I on the F trimer.

#Applications

MAGE targets the front of the antibody discovery funnel, producing a diverse candidate pool that can be triaged computationally before any protein is expressed. Because it needs only an antigen sequence, it fits situations where biological material is scarce or slow to obtain — an emerging zoonotic strain, a pathogen with no convalescent donors, or a target with no structural characterization. Groups running high-throughput B cell sequencing can also use it as an amplifier, turning existing antigen-specificity datasets into candidate repertoires for down-selection.

#Impact

MAGE demonstrated that a general protein language model, fine-tuned on antigen-linked antibody repertoires, can produce functional human antibodies against a named target without a starting template — and backed that claim with binding, neutralization, and structural data rather than in-silico metrics alone. The limitations are equally clear. The training set is heavily skewed toward coronavirus antibody-antigen pairs, and validated hit rates ranged from 28% to 45% across targets, so the model produces a shortlist rather than a finished therapeutic. The work is a preprint awaiting peer review. The Apache-2.0 repository delivers more than the paper promises — the annotated training set ships as MAGE_annotated_training_data.zip, with the full 1,000-sequence RBD output alongside it — but the checkpoint stays gated behind a Vanderbilt end-user license agreement. Its trajectory depends on data: as paired antigen-specificity datasets grow, the same recipe should reach targets well beyond the viral antigens tested here.

At a glance

Parameters
765 Million
Released
December 2024
Category
Protein
License
Apache-2.0
Organizations
Vanderbilt University Medical Center / University of Texas at Austin / Karolinska Institutet / Cleveland Clinic / National Institute of Allergy and Infectious Diseases / Griffith University / University of Washington / Vanderbilt University

Related models

  • AAMFM

    Shanghai Jiao Tong University

  • p-IgGen

    Oxford Protein Informatics Group (OPIG) / AstraZeneca

  • IgGM

    Tencent AI for Life Science Lab / Chinese Academy of Sciences / University of Chinese Academy of Sciences / Fudan University / Sun Yat-sen University / Shanghai Jiao Tong University / Guangzhou Medical University

  • Germinal

    Stanford University / Arc Institute

  • JAM

    Nabla Bio

  • SpeciefAI

    University of Edinburgh

Links

GitHub RepositoryResearch PaperHuggingFace ModelDataset

Tags

antibodyantibody_designgenerativelanguage_modeltransformer

Something wrong?

Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.