bio.rodeo
ModelsOrganizationsProvidersLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsProvidersLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
Protein

ApexAmphion

University of Pennsylvania / Chinese University of Hong Kong / Stanford University / Hangzhou Institute of Medicine, CAS

De novo antibiotic design framework coupling a 6.4B-parameter protein language model with reinforcement learning to generate antimicrobial peptides.

Released: September 2025
Parameters: 6.4 Billion

ApexAmphion is a deep-learning framework for the de novo design of peptide antibiotics, built to address the accelerating crisis of antimicrobial resistance (AMR), which is projected to cause millions of deaths annually by mid-century. Rather than screening existing compounds, ApexAmphion generates entirely new antimicrobial peptide sequences and optimizes them toward potency and drug-like properties, casting antibiotic discovery as a controllable generation problem.

The model was developed by Hanqun Cao, Marcelo D. T. Torres, Cesar de la Fuente-Nunez, and collaborators at the University of Pennsylvania, the Chinese University of Hong Kong, Stanford University, and the Hangzhou Institute of Medicine of the Chinese Academy of Sciences, and posted to bioRxiv in September 2025. It couples a 6.4-billion-parameter protein language model with reinforcement learning: the language model is first fine-tuned on curated peptide data to capture antimicrobial sequence regularities, then optimized with proximal policy optimization (PPO) against a composite reward.

That reward combines a learned minimum inhibitory concentration (MIC) classifier with differentiable physicochemical objectives, so generation is steered jointly by predicted activity and developability. Unifying generation, scoring, and multi-objective optimization in a single pipeline lets ApexAmphion produce diverse, potent candidates rapidly.

#Key Features

  • Generation-plus-RL pipeline: A large protein language model is fine-tuned on peptide data and then optimized with proximal policy optimization, unifying sequence generation, scoring, and multi-objective steering in one loop.
  • Composite reward: Optimization targets a reward that blends a learned MIC classifier with differentiable physicochemical objectives, balancing predicted potency against drug-like properties.
  • 6.4-billion-parameter backbone: The generator is built on a 6.4B-parameter protein language model, giving it broad sequence priors to draw on when proposing novel peptides.
  • High experimental hit rate: All 100 designed peptides tested in vitro showed measurable activity (100% hit rate), some with nanomolar MIC values, and 99 of 100 were broad-spectrum against at least two clinically relevant bacteria.

#Technical Details

ApexAmphion fine-tunes a 6.4-billion-parameter protein language model on curated antimicrobial peptide data, then applies reinforcement learning via PPO against a composite reward that couples a learned MIC classifier with differentiable physicochemical terms. In experimental validation, 100 designed peptides all exhibited low MIC values — reaching the nanomolar range in some cases — for a 100% in vitro hit rate, and 99 of the 100 showed broad-spectrum activity against at least two clinically relevant bacterial species. Mechanistic follow-up indicated the lead molecules kill bacteria primarily by targeting the cytoplasmic membrane. The authors describe the system as a platform for iterative steering toward potency and developability within hours. No code or model weights accompany the preprint.

#Applications

ApexAmphion is aimed at antibiotic discovery, generating candidate antimicrobial peptides for experimental testing against drug-resistant pathogens. For researchers and translational teams confronting AMR, it offers a route to rapidly propose diverse, potent leads with tunable physicochemical profiles, compressing the front end of antibiotic development from library screening to targeted in silico design followed by synthesis and assay.

#Impact

The reported 100% in vitro hit rate across 100 designs is a striking validation result for generative antibiotic design, suggesting that language-model generation paired with reinforcement learning against activity and developability rewards can yield genuinely active molecules. As with all generative antimicrobial design, dual-use and biosecurity considerations apply and warrant scrutiny. The framework is a preprint awaiting peer review, and with no released code, weights, or disclosed base model, independent reproduction and broader adoption will depend on further public release.

Citation

Preprint

DOI: 10.1101/2025.09.23.678086

Recent citations

Papers that recently cited this model.

Not enough citation data yet.

Top citations

The most-cited papers that cite this model.

Not enough citation data yet.

Where to run ApexAmphion

Providers that host ApexAmphion for inference, fine-tuning, or weight download.

No providers recorded yet. Browse all providers

Fields of citing research

Not enough data

Openness

bio.rodeo opennessClosed · low usability and reproducibility
38Closed
Usability — can I run it?29
Reproducibility — can I retrain it?37

Tags

antimicrobial_peptidesde_novo_designgenerativelanguage_modelpeptidepeptide_designprotein_designreinforcement_learningtransformer

Resources

Research Paper