Hugging Face

Download open weights for protein, DNA, RNA, single-cell, pathology and medical imaging foundation models from Hugging Face model repositories.

Website
258 models · 258 with weights

Overview

Hugging Face is the open model hub where biological foundation model weights are published, versioned, and downloaded. If you are looking for where to get the checkpoint behind a protein language model, a genome model, a single-cell foundation model, or a pathology encoder, the answer is almost always a Hugging Face repository. It is the default download surface for the bio.rodeo catalog, and it reaches domains — histopathology, radiology, biosignals — that specialized bio platforms rarely host. Each repository pairs the weights with a model card, config, and tokenizer or preprocessing code, so a download is enough to load the model locally or inside your own pipeline.

Models with weights on Hugging Face

Weights for models spanning every biological domain in the catalog live on Hugging Face, including protein language models such as ESM, ESM-3, ProstT5, and ProGen3; structure and complex predictors such as Boltz-1 and Chai-1; genome and DNA models such as Evo 2, Nucleotide Transformer, AlphaGenome, and Caduceus; single-cell foundation models such as Geneformer, scFoundation, scPRINT, and STATE; computational pathology encoders such as UNI, Virchow, CONCH, and H-optimus-0; and multimodal medical models such as MedGemma and MedSAM. RNA models such as RNA-FM and RNABERT round out the nucleic-acid side. This is a representative slice, not the full list — hundreds of catalog checkpoints resolve to a Hugging Face URL, and each model's bio.rodeo entry links directly to its repository.

Downloading weights from Hugging Face

Weights are downloaded directly from each model repository, either through the web UI, git lfs, or the huggingface_hub Python client, with support for gated access when a model requires accepting a license. Hugging Face is a weights host rather than a runtime: it does not run hosted inference or managed fine-tuning for these catalog models by default, so treat it as the place to fetch a checkpoint and then serve, fine-tune, or embed it in your own infrastructure. This makes it the natural starting point for researchers, engineers, and autonomous agents that need reproducible, directly loadable weights for offline or self-hosted use.

Download weights from Hugging Face (258)

ESM-2 & ESMFold

Meta AI

Released July 20, 2022

5.1K1.9M4.2K

Meta AI's family of protein language models scaled to 15B parameters, paired with ESMFold for fast, alignment-free atomic-level structure prediction.

Protein

Big Bird

Google Research

Released July 28, 2020

3K318.4K633

Sparse attention transformer that extends BERT to 8x longer sequences via random, local, and global attention, with genomic sequence applications.

DNA & Gene

scVI (CELLxGENE Census)

Chan Zuckerberg Initiative

Released July 1, 2024

2.5K1.7K

Variational autoencoder pretrained on 74 million human single-cell transcriptomes from the CELLxGENE Census for batch correction and cell typing.

Single-cell

LLaVA-Med

Microsoft Research

Released June 1, 2023

1.9K12.2K2.2K

Biomedical vision-language assistant for question answering on radiology and pathology images, adapted from LLaVA on PubMed Central captions.

PathologyLanguage model

UNI

Mahmood Lab

Released March 22, 2024

1.6K45.5K762

Computational pathology foundation model (ViT-L/16, DINOv2) pretrained on over 100 million H&E tiles from more than 100,000 whole-slide images.

Pathology

BioGPT

Microsoft Research Asia / Microsoft Research

Released October 19, 2022

1.5K108.3K4.5K

Generative transformer pretrained on PubMed abstracts for biomedical text generation and mining, including relation extraction and question answering.

Language model

MedSAM

Bowang Lab / University Health Network / University of Toronto / Vector Institute / Western University / New York University / Yale University

Released January 22, 2024

1.5K1.9K4.4K

Promptable foundation model for universal medical image segmentation, fine-tuned from SAM on 1.57M image-mask pairs across 10 imaging modalities.

Imaging

ProtTrans

Rostlab

Released August 1, 2021

1.4K1.3K

Suite of six protein language models, including ProtBERT and ProtT5, trained on up to 393 billion amino acids without multiple sequence alignments.

Protein

Geneformer

Broad Institute / Dana-Farber Cancer Institute

Released May 31, 2023

1.1K4.8K

Single-cell foundation model pretrained on about 30 million human transcriptomes, using rank-value encoding for context-aware gene network inference.

Single-cell

Galactica

Meta AI

Released November 16, 2022

1.1K6422.7K

Scientific large language model trained on 48 million papers, textbooks, and reference works to store, combine, and reason about scientific knowledge.

Language model

CONCH

Mahmood Lab / Brigham and Women's Hospital

Released March 19, 2024

1.1K76.3K519

Histopathology vision-language foundation model pretrained on 1.17 million image-caption pairs with contrastive and captioning objectives.

Imaging

ProteinBERT

Hebrew University of Jerusalem

Released January 13, 2022

984579

Protein language model pretrained on UniRef90 with masked language modeling and Gene Ontology annotation prediction, at 16 million parameters.

Protein

Prov-GigaPath

Microsoft Research

Released May 22, 2024

95480.2K626

Whole-slide histopathology foundation model pretrained on 1.3 billion image tiles from 171,189 clinical slides spanning 31 tissue types.

Pathology

RETFound

University College London / Google DeepMind

Released September 13, 2023

947116661

Self-supervised foundation model for retinal imaging, pretrained on 1.6 million unlabelled fundus and OCT scans to detect ocular and systemic disease.

ImagingPathology

ProtGPT2

University of Bayreuth

Released July 27, 2022

8729.2K

Autoregressive protein language model based on GPT-2 that generates de novo protein sequences sampling unexplored regions of protein space.

Protein

PLIP

Stanford University

Released August 28, 2023

83549.9K382

Vision-language foundation model for pathology, fine-tuned from CLIP on 208,414 image-text pairs for zero-shot classification and image retrieval.

Imaging

BiomedCLIP

Microsoft Research

Released March 1, 2023

670877.7K129

Biomedical vision-language model trained contrastively on 15M PubMed Central figure-caption pairs for zero-shot classification, retrieval, and VQA.

Imaging

Med-Flamingo

Stanford University / Harvard Medical School / Hospital Israelita Albert Einstein

Released July 27, 2023

620452

Multimodal medical vision-language model for few-shot visual question answering, learning new imaging tasks from in-context examples at inference.

PathologyLanguage model

MoLFormer-XL

IBM Research

Released October 3, 2022

599224.1K406

Large-scale chemical language model trained on 1.1 billion SMILES strings using linear attention transformers for molecular property prediction.

Small molecule

scFoundation

Biomap Research

Released June 6, 2024

599423

Single-cell transcriptomics foundation model with 100 million parameters, pretrained on over 50 million human scRNA-seq profiles for cell embeddings.

Single-cell

HyenaDNA

HazyResearch

Released June 27, 2023

520799

Genomic foundation model built on the Hyena operator, processing DNA at single-nucleotide resolution with context windows up to 1 million tokens.

DNA & Gene

DNABERT-2

MAGICS Lab

Released June 26, 2023

457169.7K508

Multi-species genomic foundation model swapping k-mer tokenization for byte pair encoding, matching Nucleotide Transformer with 21x fewer parameters.

DNA & Gene

Boltz-1

MIT

Released November 14, 2024

4334.1K

Open-source structure prediction model for proteins, nucleic acids, and small molecules, trained on public data to AlphaFold3-level accuracy.

Protein

BiomedGPT

Lehigh University / University of Georgia / Stanford University / Massachusetts General Hospital / University of Pennsylvania / University of Central Florida / UC Santa Cruz / UTHealth Houston / Mayo Clinic / Samsung Research America

Released August 7, 2024

4092708

Open-source, lightweight generalist vision-language foundation model for diverse biomedical imaging and text tasks.

Language modelImagingPathology

Chai-1

Chai Discovery

Released October 11, 2024

4062K

Biomolecular structure prediction foundation model covering proteins, small molecules, DNA, RNA, and glycans in a single diffusion framework.

Protein

MedSigLIP

Google Research / Google DeepMind

Released July 9, 2025

37620.4K298

Medically tuned SigLIP encoder from Google that maps medical images and text into one embedding space for zero-shot classification and retrieval.

ImagingPathology

MedGemma

Google Research / Google DeepMind

Released July 9, 2025

376112.4K1.6K

Open medical multimodal models from Google, built on Gemma 3 with a medically tuned SigLIP vision encoder for clinical text and image understanding.

Language modelImaging

MedVInT

Shanghai Jiao Tong University / Shanghai AI Laboratory

Released May 17, 2023

369236

Generative medical visual question answering model that pairs a vision encoder with a language model, trained on the 227k-pair PMC-VQA dataset.

PathologyLanguage model

CLIP-Driven Universal Model

City University of Hong Kong / Johns Hopkins University / NVIDIA

Released October 1, 2023

358677

Abdominal CT segmentation model driven by CLIP text embeddings, covering 25 organs and 6 tumor types with zero-shot extension to new categories.

Imaging

SaProt

Westlake University

Released October 1, 2023

35239.2K613

Structure-aware protein language model pairing amino acid tokens with Foldseek 3Di structural states, outperforming ESM-2 across 10 downstream tasks.

Protein

BioEmu-1

Microsoft

Released December 5, 2024

343857

Generative model that emulates protein equilibrium ensembles, sampling cryptic pockets and unfolded states far faster than molecular dynamics.

Protein

ESM-3

EvolutionaryScale

Released June 25, 2024

3142.4K2.9K

Multimodal generative protein language model reasoning jointly over protein sequence, structure, and function, trained at 98B parameters.

Protein

PubMedCLIP

Hasso Plattner Institute

Released December 27, 2021

3126.7K183

Medical-domain CLIP fine-tuned on radiology image-caption pairs from ROCO, serving as a drop-in visual encoder for medical visual question answering.

PathologyLanguage model

Evo 2

Arc Institute

Released February 21, 2025

2921.8K4K

Genomic foundation model trained on 9.3 trillion DNA base pairs across all domains of life, with 40B parameters and a 1-million-token context.

DNA & Gene

MUSK

Stanford University / Harvard Medical School

Released January 8, 2025

288241

Vision-language foundation model for precision oncology, pretrained on 50M pathology images and 1B text tokens via unified masked modeling.

PathologyLanguage model

RadFM

Shanghai Jiao Tong University / Shanghai AI Laboratory

Released August 4, 2023

267562

Radiology foundation model that reads interleaved 2D and 3D scans with text for diagnosis, visual question answering, and report generation.

ImagingLanguage model

AlphaFlow

MIT

Released February 7, 2024

260535

Protein conformational ensemble generator that fine-tunes AlphaFold 2 with flow matching, sampling protein dynamics beyond a single static structure.

Protein

Medical SAM 2

University of Oxford / National University of Singapore

Released August 1, 2024

259932

SAM2-based foundation model that segments 2D and 3D medical images by treating volumes and image sets as video object tracking.

Imaging

Borzoi

Calico Life Sciences

Released January 9, 2025

258257

Regulatory genomics model predicting cell-type-specific RNA-seq coverage from DNA sequence, unifying transcription, splicing, and polyadenylation.

DNA & Gene

RNA-FM

ml4bio / Chinese University of Hong Kong / Fudan University / Shanghai AI Laboratory

Released August 6, 2022

257386

RNA foundation model pretrained on 23.7 million non-coding RNA sequences, producing embeddings for structure prediction, annotation, and RNA design.

RNA

Evo

Arc Institute

Released November 15, 2024

2532.2K1.5K

Genomic foundation model with 7B parameters that models prokaryotic DNA, RNA, and protein at single-nucleotide resolution over a 131k-token context.

DNA & Gene

Virchow

Paige AI

Released September 14, 2023

23510.4K46

Histopathology foundation models: self-supervised vision transformers pretrained on millions of whole-slide images for tile-level feature extraction.

Pathology

Nucleotide Transformer

InstaDeep

Released January 11, 2023

2282K901

DNA foundation models from 500M to 2.5B parameters, trained on 3,200+ human genomes and 850 species for variant effect prediction.

DNA & Gene

Caduceus

Kuleshov Lab

Released March 5, 2024

2243.9K248

Bidirectional, reverse-complement equivariant DNA language models built on Mamba state space models for long-range variant effect prediction.

DNA & Gene

RhoFold+

ml4bio

Released November 1, 2024

216834243

End-to-end RNA 3D structure prediction from sequence alone, coupling the RNA-FM language model with an Invariant Point Attention structure module.

RNA

HuatuoGPT-Vision

Shenzhen Research Institute of Big Data / Chinese University of Hong Kong, Shenzhen

Released June 27, 2024

2042.1K398

Open medical multimodal LLMs (7B and 34B) for visual question answering over radiology, pathology, and endoscopy images, trained on PubMedVision.

PathologyLanguage model

Lingshu

DAMO Academy / Hupan Lab

Released June 8, 2025

201144.7K3

Generalist medical multimodal LLM for image understanding, visual question answering, and report generation across twelve-plus imaging modalities.

ImagingLanguage model

Cellpose-SAM

HHMI Janelia Research Campus

Released May 1, 2025

1962.3K

Generalist cell segmentation model pairing SAM's ViT-L encoder with Cellpose flow fields, outperforming average human annotators on its benchmark.

Imaging

MedVLM-R1

Technical University of Munich / Imperial College London / University of Oxford

Released February 26, 2025

1921K32

2B-parameter medical vision-language model that uses reinforcement learning to show interpretable reasoning for radiology visual question answering.

ImagingLanguage model

CBraMod

Zhejiang University

Released December 10, 2024

187329

EEG foundation model for brain-computer interface decoding, factorizing self-attention into parallel spatial and temporal branches.

Biosignals

SAM-Med3D

Shanghai AI Laboratory

Released October 23, 2023

185947

Fully 3D promptable segmentation foundation model for volumetric CT and MR, encoding whole volumes so anatomy can be segmented from one prompt point.

Imaging

M3D

Beijing Academy of Artificial Intelligence

Released March 31, 2024

175806454

Multimodal large language model for 3D medical imaging that handles report generation, visual question answering, and segmentation on CT volumes.

ImagingLanguage model

BiomedParse

Microsoft Research

Released November 18, 2024

171542686

Biomedical imaging foundation model that segments, detects, and recognizes structures across nine modalities from natural language prompts.

Imaging

ProtST

DeepGraphLearning

Released January 1, 2023

1697105

Multi-modal protein language model trained on sequences paired with biomedical text, enabling zero-shot function prediction and text-based retrieval.

Protein

AlphaGenome

Google DeepMind

Released June 27, 2025

1592K

DNA foundation model that predicts thousands of functional genomic tracks, from expression and splicing to chromatin, at single base-pair resolution.

DNA & Gene

MAIRA-2

Microsoft Research

Released June 6, 2024

1534.3K

Microsoft Research multimodal LLM for grounded chest X-ray report generation, localizing each described finding with bounding boxes on the image.

ImagingLanguage model

Nicheformer

Helmholtz Munich / Technical University of Munich

Released April 17, 2024

145958168

Transformer foundation model pretrained on 110M single-cell and spatial transcriptomics profiles, transferring spatial context to dissociated cells.

Single-cell

OntoProtein

Zhejiang University

Released January 28, 2022

141211152

Protein language model that fuses Gene Ontology knowledge graphs with masked language modeling, improving protein function and interaction prediction.

Protein

Med-R1

Emory University / University of Southern California / University of Tokyo / Johns Hopkins University / Georgia Institute of Technology

Released March 18, 2025

140129

Medical vision-language model trained with reinforcement learning for generalizable reasoning across eight imaging modalities and five question types.

ImagingLanguage model

Merlin

Stanford University

Released January 1, 2026

1388K454

3D vision-language foundation model for abdominal CT, pretrained on scans, radiology reports, and EHR codes for zero-shot interpretation.

ImagingLanguage model

RNABERT

Keio University

Released February 22, 2022

1345.7K56

RNA language model that learns base-level embeddings capturing sequence context and secondary structure, enabling fast structural alignment.

RNA

BrainLM

Yale University / Baylor College of Medicine / Princeton University

Released September 12, 2023

13218

fMRI foundation model pretrained with masked autoencoding on roughly 6,700 hours of recordings for clinical prediction and network discovery.

Biosignals

SCimilarity

Genentech

Released November 20, 2024

126258

Single-cell foundation model trained by metric learning to embed scRNA-seq profiles for cell type annotation and similarity search in cell atlases.

Single-cell

SegVol

Beijing Academy of Artificial Intelligence

Released November 22, 2023

124659386

Promptable 3D foundation model for volumetric CT segmentation, covering over 200 anatomical categories through point, box, and free-text prompts.

Imaging

5' UTR-LM

Princeton University

Released April 1, 2024

1222.9K95

Transformer language model for 5' UTR sequences that predicts mRNA translation efficiency, ribosome loading, and protein expression levels.

RNA

STATE

Arc Institute

Released June 27, 2025

120254624

Virtual cell transformer that predicts how cells respond to genetic, chemical, or signaling perturbations, generalizing to unseen cellular contexts.

Single-cell

HealthGPT

Zhejiang University / University of Electronic Science and Technology of China / Alibaba / Hong Kong University of Science and Technology / National University of Singapore

Released February 14, 2025

118361.6K

Medical vision-language model that unifies image comprehension and generation in one autoregressive transformer via heterogeneous LoRA adapters.

PathologyImaging

RNA-MSM

Peking University / Griffith University

Released January 11, 2024

1091.4K70

RNA language model trained on multiple sequence alignments of Rfam families, predicting secondary structure and solvent accessibility from homology.

RNA

PathAsst

Westlake University / Zhejiang University / The Ohio State University / Hangzhou City University

Released May 24, 2023

106136

Multimodal pathology assistant that answers questions about histology and cytology images, pairing the PathCLIP vision encoder with a Vicuna-13B LLM.

PathologyLanguage model

ProteinDT

UC Berkeley

Released January 1, 2025

106107

Text-guided protein design framework aligning language with sequences for text-conditioned generation, zero-shot editing, and property prediction.

Protein

NeuroLM

Shanghai Jiao Tong University / Microsoft

Released August 27, 2024

106163

Multi-task EEG foundation model that treats brain signals as a foreign language, pairing a text-aligned neural tokenizer with a GPT-2 backbone.

BiosignalsLanguage model

Neuro-GPT

University of Southern California / Université de Montréal

Released November 7, 2023

99228

EEG foundation model that pairs a convolutional encoder with a GPT backbone, pretrained by masked-segment reconstruction for low-data BCI decoding.

Biosignals

ECG-FM

University of Toronto / Vector Institute

Released August 9, 2024

98299

Open transformer foundation model for 12-lead electrocardiograms, pretrained on 1.5 million unlabeled ECGs with a wav2vec 2.0 self-supervised recipe.

Biosignals

LLaVA-Tri

UC Santa Cruz / Huazhong University of Science and Technology / Harvard University / Stanford University

Released August 6, 2024

964411

Medical vision-language model trained on the MedTrinity-25M dataset, answering questions and generating text about radiology and histology images.

Language modelPathology

ProLLaMA

PKU-YuanGroup

Released February 26, 2024

95325207

Protein large language model adapted from LLaMA-2 that unifies sequence generation and superfamily classification in one 7B-parameter framework.

Protein

Brant

Zhejiang University

Released December 10, 2023

9542

500M-parameter transformer model pretrained on intracranial SEEG recordings for neural signal forecasting, imputation, and seizure detection.

Biosignals

Hibou

HistAI

Released June 7, 2024

9236.6K79

Histopathology foundation models pretrained with DINOv2 on over 1 million whole-slide images, released as Hibou-B and Hibou-L under Apache 2.0.

Pathology

GPN-MSA

UC Berkeley

Released October 11, 2023

90193349

DNA language model for variant effect prediction across coding and non-coding regions, using whole-genome alignments of 100 vertebrate species.

DNA & Gene

Med-MoE

Zhejiang University / National University of Singapore / Peking University

Released April 16, 2024

90157

Lightweight mixture-of-experts medical vision-language model routing visual question answering and image classification to domain-specific experts.

ImagingLanguage modelPathology

OpenPhenom-S/16

Recursion Pharmaceuticals

Released November 12, 2024

889.4K78

Cell Painting microscopy foundation model, a channel-agnostic masked autoencoder producing morphological embeddings for zero-shot phenotypic analysis.

Imaging

Cell2Sentence

Yale University

Released July 21, 2024

86869873

Framework turning single-cell expression profiles into ranked gene-name sequences, letting off-the-shelf language models generate and annotate cells.

Single-cell

Qilin-Med-VL

Alibaba Group

Released October 27, 2023

82665

Chinese medical vision-language model pairing a Vision Transformer with an LLM to caption medical images and answer clinical questions in Chinese.

Language modelPathology

CXR Foundation

Google Research

Released August 2, 2024

82133199

Chest X-ray embedding model built on ELIXR, producing image and image-text embeddings for data-efficient and zero-shot radiograph classification.

Imaging

VISTA3D

NVIDIA

Released June 7, 2024

819.7K292

Medical image segmentation foundation model for 3D CT and MRI, covering 127 anatomical classes automatically plus interactive point-prompt refinement.

Imaging

BoltzGen

MIT

Released November 24, 2025

811K

All-atom generative model for de novo protein and peptide binder design against diverse biomolecular targets, wet-lab validated across 26 targets.

ProteinSmall molecule

CellFM

Sun Yat-sen University

Released June 6, 2024

79110

Single-cell foundation model with 800M parameters trained on ~100 million human cells, for annotation, perturbation prediction, and gene analysis.

Single-cell

CheXagent

Stanford University

Released January 22, 2024

77908230

Instruction-tuned vision-language foundation model for chest X-ray interpretation, with 8 billion parameters spanning eight clinical task types.

ImagingLanguage model

LLaVA-Rad

Microsoft Research

Released February 20, 2025

7460658

Chest X-ray vision-language model that drafts the findings section of a radiology report, at 7B parameters small enough to run on a single GPU.

ImagingLanguage model

Ankh

Technical University of Munich

Released January 16, 2023

733.3K249

Parameter-efficient protein language model that matches larger models such as ESM-2 on protein prediction tasks using under 10% of the parameters.

Protein

TxGemma

Google DeepMind / Google Research

Released March 25, 2025

6610835

Open therapeutics foundation models from Google, built on Gemma-2, for drug-discovery property prediction and conversational reasoning.

Language modelSmall molecule

tGPT

Tianjin Medical University Cancer Institute and Hospital

Released April 20, 2023

6219917

Single-cell foundation model pre-trained on 22 million transcriptomes, using rank-based gene encoding for clustering and trajectory inference.

Single-cell

GPFM

Hong Kong University of Science and Technology / Sun Yat-sen University / Southern Medical University / Chinese University of Hong Kong

Released November 1, 2025

56129

Histopathology foundation model extracting general-purpose features from H&E patches by distilling the UNI, Phikon, and CONCH pathology encoders.

Pathology

scPRINT

Institut Pasteur

Released July 29, 2024

55156

Single-cell foundation model pre-trained on 50 million cells for gene network inference, denoising, and cell type prediction.

Single-cell

scPRINT

Institut Pasteur / CNRS

Released April 16, 2025

55156

Single-cell foundation model pre-trained on 50 million cells that infers cell-specific gene regulatory networks from transformer attention matrices.

Single-cell

ProGen3

Profluent

Released April 16, 2025

54264114

Sparse mixture-of-experts autoregressive protein language model family pretrained on 1.5 trillion amino acid tokens with compute-optimal scaling.

Protein

Species-Aware DNA Language Model

Technical University of Munich

Released January 27, 2023

537.8K18

Masked DNA language model trained on 800+ species with explicit species conditioning, separating conserved regulatory motifs from background bias.

DNA & Gene

xTrimoPGLM

BioMap / Tsinghua University

Released January 11, 2024

5321

Unified 100-billion-parameter protein language model combining autoencoding and autoregressive objectives for protein understanding and generation.

Protein

Species-Aware DNA LM

Technical University of Munich

Released January 27, 2023

537.8K29

Masked DNA language model trained on over 800 vertebrate genomes and conditioned on species identity to learn conserved regulatory sequence features.

DNA & Gene

DNABERT-S

MAGICS Lab

Released February 13, 2024

5324K130

DNA embedding model built on DNABERT-2, using contrastive learning to cluster sequences by species for metagenomic binning without labeled data.

DNA & Gene

GenerRNA

Preferred Networks

Released October 1, 2024

4719

Transformer-based generative language model for de novo RNA design, pretrained on 16 million non-coding RNA sequences from RNAcentral.

RNA

HeartLang

Peking University

Released February 15, 2025

4757

ECG foundation model that treats heartbeats as words and rhythm strips as sentences, using heartbeat-level tokenization for diagnostic classification.

Biosignals

OPERA

University of Cambridge

Released June 23, 2024

4683

Respiratory acoustic foundation models pretrained on roughly 136K cough and breathing recordings for disease detection and lung function estimation.

Biosignals

EndoChat

Chinese University of Hong Kong / Huawei / Technical University of Munich / University of Strasbourg / Shandong University / Chinese Academy of Sciences

Released January 20, 2025

442651

Grounded multimodal language model for endoscopic surgery, supporting visual dialogue, region-based question answering, and bounding-box grounding.

ImagingLanguage model

ERNIE-RNA

Tsinghua University

Released March 17, 2024

431.8K44

RNA language model that builds base-pairing constraints into self-attention, pretrained on 20.4 million sequences for structure and function tasks.

RNA

GEM (Grounded ECG understanding with Multimodal LLM)

National University of Singapore / Peking University

Released March 8, 2025

42209192

Multimodal LLM unifying 12-lead ECG time series, ECG images, and text for grounded, clinician-aligned electrocardiogram interpretation.

BiosignalsLanguage model

MedPLIB

Baidu / China Agricultural University / Chinese Academy of Sciences / Peking University

Released December 12, 2024

4016134

Biomedical multimodal LLM that answers questions about medical images and returns pixel-level segmentation masks, using a mixture-of-experts design.

ImagingLanguage model

QoQ-Med

MIT

Released May 31, 2025

4052552

Multimodal clinical foundation model reasoning jointly over 2D and 3D medical images, ECG time-series, and text reports across nine clinical domains.

ImagingBiosignalsLanguage model

Spark3D (S3D)

German Cancer Research Center (DKFZ) / Heidelberg University / Helmholtz Imaging / National Center for Tumor Diseases (NCT) Heidelberg / FLOY / Humanitas University

Released October 30, 2024

4037157

Masked-autoencoder foundation model that pre-trains a 3D Residual Encoder U-Net on roughly 39,000 brain MRIs for volumetric image segmentation.

Imaging

CXR-LLaVA

Seoul National University / Gwangju Institute of Science and Technology

Released October 22, 2023

3913354

Chest X-ray vision-language model that generates free-text radiology reports, pairing a CXR-specific image encoder with a 7B LLaMA-2 language model.

ImagingLanguage model

Compute-Optimal PLM

BioMap

Released June 9, 2024

3811

Scaling-law study of protein language models identifying compute-optimal training for causal and masked objectives on 939 million protein sequences.

Protein

NatureLM

Microsoft Research AI for Science

Released February 11, 2025

3546

Unified science foundation model treating molecules, proteins, RNA, DNA, and materials as one sequence language, in 1B, 8B, and 46.7B sizes.

Language modelSmall moleculeProtein

HeAR (Health Acoustic Representations)

Google Research

Released December 4, 2023

3488123

Health acoustics foundation model that turns short clips of coughs and breaths into embeddings for building acoustic biomarker models with less data.

Biosignals

ECGFounder

Peking University / Harvard Medical School / Emory University

Released October 5, 2024

34124142

Convolutional ECG foundation model trained on expert annotations spanning 150 diagnostic categories, with 12-lead and single-lead wearable variants.

Biosignals

NormWear

University of California, San Diego

Released December 12, 2024

3212059

Multimodal foundation model for wearable physiological sensing across PPG, ECG, EEG, GSR, and IMU signals, using channel-aware attention.

Biosignals

CellViT

Institute for AI in Medicine

Released June 14, 2023

31391

Vision Transformer for cell instance segmentation and classification in H&E whole-slide images, extended by CellViT++ with foundation backbones.

Imaging

PULSE

The Ohio State University / Carnegie Mellon University

Released October 21, 2024

301.5K67

Multimodal large language model that interprets 12-lead electrocardiogram images, answering open-ended clinical questions and generating ECG reports.

BiosignalsImaging

Pinal

Westlake University

Released April 2, 2025

291694

De novo protein design from natural language: a 16B-parameter framework turning text descriptions into sequences via structure-conditioned generation.

Protein

MedRegA

Hong Kong University of Science and Technology / Sun Yat-sen University

Released October 24, 2024

291147

Region-aware bilingual medical multimodal LLM that handles image- and region-level vision-language tasks across eight imaging modalities.

PathologyLanguage model

MedDr

Hong Kong University of Science and Technology

Released April 23, 2024

2946100

Generalist medical vision-language foundation model with 40B parameters, spanning radiology, pathology, dermatology, retinography, and endoscopy.

ImagingLanguage model

BrainOmni

Tsinghua University / Shanghai AI Laboratory / University of Cambridge / University College London

Released May 18, 2025

2871

Brain foundation model unifying EEG and MEG in a single encoder via a shared discrete tokenizer that transfers across sensor layouts and montages.

Biosignals

AIDO.RNA

genbio.ai

Released November 28, 2024

271.2K169

RNA foundation model with 1.6 billion parameters, pretrained on 42 million non-coding RNA sequences for structure prediction and RNA sequence design.

RNA

MoME

Beijing Institute of Technology / Imperial College London / Beijing Tiantan Hospital / Capital Medical University

Released May 16, 2024

2631

Universal brain lesion segmentation for multi-modal brain MRI, using a Mixture of Modality Experts to span diverse modalities and lesion types.

Imaging

GenomeOcean

DOE Joint Genome Institute / Northwestern University / Johns Hopkins University / University of California, Merced / University of California, Berkeley / Miami University / Illumina

Released February 5, 2025

24855150

4B-parameter generative genome foundation model trained on assembled environmental metagenomes for microbial representation and de novo DNA design.

DNA & Gene

ProTrek

Westlake University

Released June 1, 2024

2449211

Tri-modal protein language model aligning sequence, structure, and text in one embedding space for natural-language search over billions of proteins.

Protein

Evolla

Westlake University

Released January 6, 2025

231169

Multimodal 80B-parameter protein-language model that answers natural language questions about protein function from sequence and structure.

Protein

AIDO.Protein

genbio.ai

Released November 29, 2024

2399169

Mixture-of-experts protein language model scaling to 16 billion parameters, applied to variant effect prediction and de novo protein design.

Protein

Proteina-Complexa

NVIDIA

Released March 16, 2026

22128401

Flow-matching generative model for de novo atomistic protein binder design against protein and small-molecule targets, including carbohydrate binders.

Protein

FetalCLIP

Mohamed bin Zayed University of Artificial Intelligence / Corniche Hospital

Released February 20, 2025

2270

Vision-language foundation model for fetal ultrasound, pretrained on 210,035 image-text pairs for plane classification, biometry, and segmentation.

Imaging

BiMediX2

Mohamed bin Zayed University of Artificial Intelligence

Released December 10, 2024

212174

Bilingual Arabic-English medical multimodal model built on Llama 3.1 for radiology, CT, and histology image understanding and question answering.

Language modelImagingPathology

MELP

The University of Hong Kong

Released June 27, 2025

191631

Multi-scale ECG-language model that aligns 12-lead ECG signals with clinical text at token, beat, and rhythm levels for zero-shot cardiac diagnosis.

BiosignalsLanguage model

AIDO.Cell

genbio.ai

Released November 28, 2024

18112169

Single-cell RNA-seq foundation model pretrained on 50 million human cells, encoding the full transcriptome for annotation and perturbation modeling.

Single-cell

Orthrus

Bowang Lab

Released October 12, 2024

1821.3K128

Mamba-based mature RNA foundation model, contrastively trained on splice isoforms and 400+ mammalian species orthologs for mRNA property prediction.

RNA

AIDO.DNA

genbio.ai

Released December 1, 2024

1788169

DNA foundation model scaling an encoder-only transformer to 7 billion parameters for variant effect prediction, gene expression, and sequence design.

DNA & Gene

LUNA

ETH Zurich

Released October 25, 2025

176.3K131

EEG foundation model whose learned queries map any electrode montage into a fixed latent space, scaling linearly in the number of channels.

Biosignals

Tahoe-x1

Tahoe Therapeutics

Released October 23, 2025

1536160

Perturbation-trained single-cell foundation models (up to 3B parameters) that jointly model genes, cells, and compounds for precision oncology tasks.

Single-cellSmall molecule

D-BETA

Singapore Management University / Eindhoven University of Technology

Released October 3, 2024

1410136

ECG foundation model pretrained on 12-lead waveforms paired with clinical reports, enabling label-efficient and zero-shot cardiac diagnosis.

BiosignalsLanguage model

Chiron-o1

Shanghai AI Laboratory / Fudan University / Shanghai Jiao Tong University

Released June 20, 2025

142060

Medical multimodal LLM (2B and 8B) trained for generalizable, step-by-step clinical reasoning via Mentor-Intern Collaborative Search.

PathologyImaging

ESMBind & QBind

Independent Researcher

Released November 14, 2023

1448

LoRA and QLoRA fine-tuning of ESM-2 for token-level prediction of protein binding sites and post-translational modification sites from sequence alone.

Protein

PLAID

UC Berkeley / Genentech

Released December 2, 2024

14127

Latent diffusion model for controllable all-atom protein generation that co-designs sequence and structure while training on sequences alone.

Protein

Path Foundation

Google Research

Released December 19, 2023

13106

Histopathology foundation model that encodes 224x224 H&E patches into compact 384-dimensional embeddings for tumor and biomarker classifiers.

Pathology

STACK

Arc Institute / Stanford University

Released January 9, 2026

13142

Single-cell foundation model using tabular attention over context cells to predict responses to arbitrary perturbations without fine-tuning.

Single-cell

ProCyon

Harvard Medical School / Kempner Institute

Released December 11, 2024

1360

Multimodal foundation model integrating protein sequence, structure, and natural language to model and generate protein phenotypes across scales.

ProteinLanguage modelSmall molecule

ESMC

Biohub

Released May 27, 2026

122.2M2.9K

Protein language model trained on roughly 2.8 billion sequences, forming the representation core of Biohub's world model of protein biology.

Protein

Scooby

Technical University of Munich / Helmholtz Munich / Harvard Medical School / Broad Institute / Harvard University

Released October 1, 2025

1226569

Predicts single-cell scRNA-seq coverage and scATAC-seq insertion profiles from DNA sequence, adapting the Borzoi trunk with a cell-specific decoder.

Single-cell

ESMFold2

Biohub

Released May 27, 2026

12370.1K2.9K

Structure-prediction and design engine that turns ESMC sequence representations into all-atom 3D structures of proteins and biomolecular complexes.

Protein

E1

Profluent

Released November 13, 2025

1115.1K114

Retrieval-augmented protein encoders that fuse homologous sequences into a single-pass transformer for variant effect and contact prediction.

Protein

MULAN

Skolkovo Institute of Science and Technology

Released May 30, 2024

106125

Multimodal protein language model extending ESM-2 and SaProt with a Structure Adapter over residue torsion angles for protein function prediction.

Protein

UniBiomed

Hong Kong University of Science and Technology / Weill Cornell Medicine / Harvard University

Released April 30, 2025

1019672

Universal foundation model that jointly generates diagnostic text and segments the corresponding targets across ten biomedical imaging modalities.

ImagingLanguage model

MuLan-Methyl

University of Tübingen

Released July 25, 2023

1067

Multi-language transformer framework using five pre-trained language models to predict DNA methylation (6mA, 4mC, 5hmC) across species.

DNA & Gene

SigPhi-Med

Chongqing University of Technology

Released July 1, 2025

945

Biomedical vision-language assistant for medical visual question answering, pairing Phi-2 with a vision encoder in a 4.2B-parameter model.

ImagingLanguage model

BrainFM

Johns Hopkins University / Massachusetts General Hospital / Harvard Medical School / Danish Research Centre for Magnetic Resonance / University College London

Released August 30, 2025

920

Modality-agnostic foundation model for human brain imaging that runs five core neuroimaging tasks across uncalibrated CT and MRI without retraining.

Imaging

MAMMAL

IBM Research

Released October 28, 2024

91.1K120

Multi-modal, multi-task biological foundation model trained on 2 billion samples spanning proteins, small molecules, and single-cell gene expression.

ProteinSmall moleculeSingle-cell

scConcept

Theis Lab / Helmholtz Munich

Released October 14, 2025

88737

Single-cell foundation model learning technology-agnostic cell embeddings by contrasting cell views rather than reconstructing gene expression counts.

Single-cell

BioMed Multi-View

IBM Research

Released October 25, 2024

86.3K46

Molecular foundation model that late-fuses graph, image, and SMILES encoders into one embedding for molecular property and drug target prediction.

Small molecule

SeqDance / ESMDance

Columbia University

Released October 11, 2024

86861

Protein language models trained on biophysical dynamics from MD simulations and normal-mode analysis; ESMDance builds on ESM2 for variant effects.

Protein

X-Cell

Xaira Therapeutics

Released March 17, 2026

8106

Diffusion language model with 4.9 billion parameters that predicts genome-wide CRISPRi perturbation responses in single-cell transcriptomes.

Single-cell

DNABERT

Northwestern University

Released February 4, 2021

812.7K772

Bidirectional transformer for DNA using k-mer tokenization, fine-tunable for promoter, splice site, and transcription factor binding prediction.

DNA & Gene

CryoFM

ByteDance Seed

Released October 11, 2024

72135

Generative foundation model for cryo-EM density maps using flow matching, enabling zero-shot denoising, map sharpening, and missing wedge restoration.

Imaging

ChatCell

ZJUNlp

Released February 13, 2024

6951

Conversational T5-based framework that turns scRNA-seq data into cell sentences for cell type annotation and drug sensitivity prediction.

Single-cellLanguage model

H-optimus-0

Bioptimus

Released July 11, 2024

642.9K109

Histopathology vision transformer with 1.1B parameters, pretrained on patches from 500,000 H&E whole-slide images across 4,000 clinical practices.

Pathology

Proteo-R1

Stanford University / University of Tokyo / RIKEN Center for Advanced Intelligence Project / Chinese University of Hong Kong

Released May 1, 2026

53.2K64

Reasoning-guided foundation model for de novo antibody CDR design, pairing a multimodal LLM understanding expert with a Boltz-1 diffusion expert.

Protein

LucaVirus

Sun Yat-sen University / University of Sydney

Released June 14, 2025

515976

Multimodal viral foundation model over nucleotide and protein sequence, built for virus discovery, function annotation, and antibody design.

DNA & GeneProtein

PINNACLE

Harvard University

Released August 1, 2024

5109

Geometric deep learning model generating context-aware protein representations across 156 cell-type contexts from a multi-organ single-cell atlas.

Single-cell

GenBio-PathFM

genbio.ai

Released March 20, 2026

477037

Histopathology foundation model with 1.1B parameters, trained entirely on public data using JEDI, a dual-stage strategy combining JEPA and DINO.

Pathology

PULSAR

Stanford University

Released November 26, 2025

416636

Hierarchical single-cell foundation model that turns scRNA-seq profiles into zero-shot donor-level embeddings for disease and biomarker prediction.

Single-cellProtein

moPPIt

Duke University

Released July 31, 2024

414

De novo peptide binder design framework that targets specific motifs, including disordered regions and conserved epitopes, from target sequence alone.

Protein

NeuroVFM

University of Michigan / University of Cologne

Released November 23, 2025

471759

Generalist neuroimaging vision foundation model pretrained on 5.24M clinical MRI and CT volumes for radiologic diagnosis and report generation.

Imaging

Derm Foundation

Google Research

Released December 19, 2023

443137

Google's dermatology image embedding model that produces 6144-dimensional embeddings for data-efficient skin-condition classifiers.

Imaging

TEA

Biozentrum / University of Basel / SIB Swiss Institute of Bioinformatics

Released November 27, 2025

43.6K24

Protein sequence encoder that maps ESM2 embeddings to a learned 20-letter alphabet for structure-quality remote homology detection at MMseqs2 speed.

Protein

SpaFoundation

Central South University

Released August 11, 2025

38

Histology vision transformer with 80M parameters that predicts spatial gene expression from H&E tissue images and transfers to tumor detection.

PathologySpatial omics

structRFM

University of Science and Technology of China

Released August 7, 2025

34436

RNA foundation model pretrained jointly on sequences and secondary structures for structure prediction, homology and splice site classification.

RNA

gRNAde

MRC Laboratory of Molecular Biology / University of Cambridge

Released December 1, 2025

320312

RNA inverse-folding model that generates sequences predicted to fold into a target 3D backbone, capturing non-canonical pairs and tertiary motifs.

RNA

DISCO

FutureHouse / Mila / McGill University

Released April 6, 2026

3210

Multimodal diffusion model that co-designs protein sequence and 3D structure around cofactors and small molecules for de novo heme enzyme design.

Protein

Lingshu-Cell

DAMO Academy

Released March 26, 2026

3

Virtual cell model using masked discrete diffusion over the whole transcriptome to simulate scRNA-seq perturbation responses across tissues.

Single-cell

GPN

Song Lab

Released October 31, 2023

31.9K349

DNA language model for genome-wide variant effect prediction, trained by masked language modeling on multispecies genomes with no labeled data.

DNA & Gene

CortexMAE

Sophont / MedARC

Released October 15, 2025

39153

fMRI foundation model trained on cortical flat-map videos with masked autoencoding, showing power-law scaling on brain activity reconstruction.

Biosignals

ProteomeLM

EPFL

Released August 1, 2025

323736

Proteome-scale protein language model whose representations enable zero-shot protein-protein interaction and gene essentiality prediction.

Protein

OmniNA

Beijing Institute of Genomics / Chinese Academy of Sciences

Released April 13, 2026

3106

Generative DNA foundation model trained on 91.7M nucleotide sequences and annotations for species classification and mutation effect prediction.

DNA & Gene

ProtProfileMD

Helmholtz Munich / Rostlab / Seoul National University

Released January 22, 2026

336

LoRA adapter on ProstT5 predicting per-residue distributions over Foldseek 3Di tokens, capturing conformational flexibility from MD trajectories.

Protein

PeptideCLM-2

University of Texas at Austin / Novo Nordisk

Released April 17, 2026

210

Chemical language models pretrained on SMILES for therapeutic peptides, natively representing non-canonical residues, cyclization, and conjugation.

Small moleculeProtein

CellHermes

Tongji University / Helmholtz Munich

Released November 28, 2025

210130

Single-cell foundation model adapting LLaMA-3.1-8B with LoRA, recasting transcriptomes and protein interaction networks as natural-language Q&A pairs.

Single-cellRNA

Reverse Distillation (ESM-2)

Duke University

Released March 8, 2026

2114

Post-hoc method that restores monotonic scaling to ESM-2 embeddings, yielding Matryoshka-style nested representations for variant effect prediction.

Protein

Atacformer

University of Virginia

Released November 4, 2025

29728

Transformer foundation model for single-cell ATAC-seq that embeds both cells and cis-regulatory elements for annotation and batch correction.

Single-cellDNA & Gene

TD3B

Duke University

Released May 15, 2026

2

Sequence-based discrete-diffusion framework that designs peptide binders with specified agonist or antagonist behavior against GPCR targets.

Protein

EXAONE Path 2.5

LG AI Research

Released December 16, 2025

21525

Pathology foundation model that aligns whole-slide images with genomic, epigenetic, and transcriptomic data for patient-level tumor representations.

PathologySpatial omics

GCP-VQVAE

University of Missouri

Released October 1, 2025

243

Protein structure tokenizer that maps 3D backbones to discrete tokens with an SE(3)-equivariant encoder preserving orientation and chirality.

Protein

ProFam

University College London / Technical University of Munich

Released December 21, 2025

258

Protein-family language model trained on unaligned homolog sets for zero-shot variant fitness prediction and design. ProFam-1 holds 251M parameters.

Protein

mRNA-GPT

Chinese Academy of Sciences

Released April 2, 2026

24

Autoregressive model for therapeutic mRNA design that jointly generates 5' UTR, CDS, and 3' UTR, pretrained on 30 million full-length natural mRNAs.

RNA

FlashPPI

Tatta Bio

Released March 1, 2026

140.8K40

Contrastive model built on a genomic language model that predicts physical protein-protein interactions across a microbial proteome in linear time.

Protein

LinkLlama

UC Berkeley

Released April 16, 2026

11410

Molecular linker design model fine-tuned from Llama 3 that emits PROTAC and fragment linkers as SMILES from natural-language geometry prompts.

Small molecule

MMAI Liquid Foundation Model (LFM2-2.6B-MMAI)

Liquid AI / Insilico Medicine

Released March 3, 2026

16.7K

Small-molecule drug discovery foundation model covering ADMET, retrosynthesis, drug-target activity, and molecular optimization in a 2.6B checkpoint.

Small moleculeLanguage model

Suiren-1.0

Golab (SAIS Physics Lab)

Released March 23, 2026

117

Molecular foundation models pretrained on density functional theory data, encoding 3D geometry and quantum behavior for ADMET and drug discovery.

Small molecule

UltraNMR

Hong Kong University of Science and Technology / Hunan University / Institute of Materia Medica, CAMS & PUMC / Xiamen University / Shanghai AI Laboratory

Released June 18, 2026

11

NMR foundation model trained on 158 million simulated 1H and 13C spectra, transferring simulation-learned representations to real experimental data.

Small moleculeMetabolomics

EVA

GENTEL Lab

Released March 24, 2026

182

Generative RNA foundation model trained on 114 million full-length sequences for de novo design of tRNAs, aptamers, CRISPR guide RNAs, and mRNAs.

RNA

CodonTranslator

University of Maryland, College Park

Released November 24, 2025

15

Conditional codon language model with 150M parameters that generates species-optimized coding sequences from a protein and its taxonomic lineage.

DNA & GeneRNA

vir2vec

University of Florida

Released December 12, 2025

12743

Pan-viral genomic language model producing fixed genome-level embeddings of viral DNA and RNA, reused across classification tasks without retraining.

DNA & Gene

MiAE (Masked Invariant Autoencoders)

ETH Zurich

Released May 18, 2026

11118

SE(3)-invariant masked autoencoder that learns protein fold representations from AlphaFold-DB structures, supporting zero-shot fold classification.

Protein

D3LM

Renmin University of China

Released March 2, 2026

163

DNA foundation model using masked discrete diffusion to unify bidirectional sequence understanding and de novo generation in one architecture.

DNA & Gene

CryoSiam

European Molecular Biology Laboratory

Released November 12, 2025

120

Self-supervised Siamese network for cryo-electron tomography, enabling zero-shot denoising, segmentation, and macromolecule detection in tomograms.

Imaging

GENERator-v2

Beijing Zhongguancun Academy / Mila / Université de Montréal / University of Science and Technology of China / HEC Montréal

Released January 29, 2026

1461

Family of autoregressive genomic foundation models that reconcile k-mer tokenization with single-nucleotide resolution at contexts up to 98k bp.

DNA & Gene

BOTANIC-0

Living Models

Released February 23, 2026

1180

Plant genomic foundation models from 0.1B to 1B parameters, pretrained on 43 phylogenetically diverse plant genomes for variant effect prediction.

DNA & Gene

RigidSSL

Chinese University of Hong Kong

Released March 2, 2026

120

Self-supervised SE(3) geometric pretraining for protein backbone generators, improving designability, motif scaffolding, and conformational ensembles.

Protein

yakRNA Design

Stanford University

Released April 24, 2026

4

110M-parameter RNA language model that designs sequences from secondary structure, motif, and Gene Ontology constraints via discrete diffusion.

RNA

SciCore-Omics

Nanjing University / OpenBMB / Tsinghua University

Released May 30, 2026

8110

Tri-modal foundation model unifying histology images, spatial transcriptomics, and language for zero-shot pathology and spatial biology reasoning.

PathologySpatial omics

muat

University of Helsinki

Released April 3, 2026

8

Transformer that classifies tumour types and subtypes from somatic variants in whole-genome and whole-exome data, with auto-downloading checkpoints.

DNA & Gene

FishMamba-1

Institute of Hydrobiology, Chinese Academy of Sciences

Released March 9, 2026

13

Genomic foundation model for Cypriniformes fish, built on a Mamba-2 state space model with a 32 kb context window for long-range genome modeling.

DNA & Gene

EVA

Scienta Lab

Released February 10, 2026

96

Cross-species multimodal foundation model of immunology and inflammation, harmonizing transcriptomics and histology into patient-level embeddings.

Single-cellRNAPathology

UBio-MolFM

IQuestLab

Released February 13, 2026

633

Universal all-atom machine-learning force field for molecular dynamics, with ab initio-level accuracy on solvated biomolecules of ~1,500 atoms.

Small moleculeProtein

Emap2lig

Kihara Lab / Purdue University

Released June 4, 2026

2

Cryo-EM ligand modeling pipeline that detects bound ligand densities in a map, then reconstructs their atomic structures with a diffusion model.

ImagingSmall molecule

STMDiT

ETH Zurich / University Hospital Basel

Released May 29, 2026

Diffusion transformer for virtual tissue synthesis, generating H&E histopathology patches conditioned on spatial gene expression and morphology.

PathologySpatial omics

PlasmidLM

University College London

Released May 19, 2026

2

Promptable DNA language model that generates multi-kilobase plasmid sequences from plain-language component specs, refined with verifiable rewards.

DNA & Gene

eccDNAMamba

Brown University

Released November 25, 2025

5

Bidirectional state-space (Mamba-2) genomic model for ultra-long extrachromosomal circular DNA, scaling linearly with sequence length.

DNA & Gene

DNAGPT2

CEITEC Masaryk University

Released June 12, 2026

Family of ten compact GPT-2 decoder-only DNA language models spanning BPE vocabularies from 16 to 8192 tokens, built for lossless genome compression.

DNA & Gene

DecoderTCR

Biohub / University of Chicago

Released February 4, 2026

8

Masked language model for T-cell receptor and peptide-MHC binding prediction, with compositional pretraining and non-autoregressive decoding.

Protein

PlantBiMoE

Huazhong University of Science and Technology

Released December 8, 2025

48

Plant genome foundation model pairing a bidirectional Mamba backbone with sparse Mixture-of-Experts, pretrained on 25.4B nucleotides from 42 species.

DNA & Gene

Cryo-IEF

Westlake University

Released November 6, 2024

72

Cryo-EM foundation model pre-trained on 65 million particle images, enabling zero-shot classification, pose clustering, and quality assessment.

Imaging

PepForge

Technical University of Berlin

Released June 2, 2026

4

Generative model for chemically modified and macrocyclic peptides that builds molecules in HELM notation, supporting de novo design and infilling.

ProteinSmall molecule

PerturbGen

Wellcome Sanger Institute

Released March 5, 2026

25

Generative single-cell foundation model trained on 100M+ transcriptomes that predicts how genetic perturbations reshape cell trajectories over time.

Single-cell

CDS-BART

MOGAM Institute for Biomedical Research

Released March 12, 2026

7

Coding-sequence foundation model for mRNA design, pretrained as a BART denoising encoder-decoder on mRNA from nine taxonomic groups.

RNA

ProtSent

Hebrew University of Jerusalem / Ben-Gurion University of the Negev

Released May 7, 2026

137

Protein sequence embedding model, contrastively fine-tuned from ESM-2, that places functionally and structurally related proteins close together.

Protein

Proust

ETH Zurich

Released February 2, 2026

9

Causal 309M-parameter protein language model that scores variant fitness zero-shot and generates sequences, reaching 0.390 Spearman on ProteinGym.

Protein

Carbon

Hugging Face / Beijing Zhongguancun Academy / TIGEM

Released May 1, 2026

6.2K201

Autoregressive DNA foundation model for variant effect prediction, using 6-mer tokenization to match Evo2-7B win rates at far higher throughput.

DNA & Gene

ConvergeCELL

Converge Bio

Released May 7, 2026

29

Virtual cell foundation model pretrained on over 23 million cells from 5,000 patient samples for drug target and biomarker discovery.

Single-cell

HELM-BERT

Kyoto University

Released December 29, 2025

47514

Peptide language model trained on HELM notation, a DeBERTa encoder for property prediction on macrocyclic and non-canonical medium-sized peptides.

Small molecule

Halo

Duke University School of Medicine

Released April 6, 2026

Whole-cell segmentation model for spatial transcriptomics that fuses DAPI nuclear images with RNA transcript density to recover true cell boundaries.

Spatial omics

Stoic

University of Basel

Released March 16, 2026

14815

Predicts protein complex stoichiometry from amino acid sequence alone, ranking copy numbers in seconds and exporting AlphaFold3-ready JSON files.

Protein

sm_protgpt2

University of Naples Federico II / University of Bern

Released May 5, 2026

8

Three fixed ProtGPT2 fine-tunes specialized for metalloprotein generation, trained on ProteinMPNN-derived synthetic sequences.

Protein

GlycanGT

Nagoya University

Released December 16, 2025

3

Graph transformer foundation model for glycans, learning reusable embeddings of branched carbohydrate structures for glycomics prediction tasks.

Small molecule

GermRL

Johns Hopkins University

Released June 11, 2026

51

Reinforcement learning framework that fine-tunes the ProGen2-OAS antibody language model with GRPO to cut germline bias in generated sequences.

Protein

Gengram

Zhejiang Lab

Released January 29, 2026

51

Retrieval-augmented genomic foundation model that gives transformer backbones a hash-based k-mer motif memory for functional genomics tasks.

DNA & Gene

GenBloom

Helmholtz Munich / LMU Munich

Released May 28, 2026

3

Genetically aligned foundation model for blood smear cytology that links single-cell morphology to the chromosomal aberrations behind AML and APL.

Pathology

C3P

University of Toronto

Released May 24, 2026

1

Contrastive promoter-protein pretraining that aligns bacterial promoters with their encoded proteins to learn regulatory genomics representations.

DNA & Gene

MOJO

InstaDeep

Released June 25, 2025

9840

Bimodal masked language model that jointly encodes bulk RNA-seq expression and DNA methylation into patient-level embeddings for cancer genomics.

RNADNA & Gene

MolDeBERTa

Florida International University

Released February 17, 2026

55

SMILES molecular encoder on a DeBERTaV2 backbone, pretrained on 123M PubChem molecules with physicochemical and structural-similarity objectives.

Small molecule

Molexar

Peking University

Released June 24, 2026

147

Multimodal molecular generation model for drug design, conditioned on properties, pharmacophores, protein sequences, or protein binding pockets.

Small moleculeProtein

PlantGeneAnn

Huazhong Agricultural University

Released June 25, 2026

12113

Plant genome foundation model for ab initio gene structure annotation, predicting genes, coding sequences, and exons at single-nucleotide resolution.

DNA & Gene

MicroGenomer

BGI Research

Released December 29, 2025

10

470M-parameter microbial genome foundation model trained on 234.5B base pairs for multi-scale genomic representation and trait prediction.

DNA & Gene

LucaPhylo

Alibaba Cloud / Sun Yat-sen University / University of Sydney

Released May 26, 2026

13

Hyperbolic protein language model for alignment-free phylogenetic inference, turning ESM2-650M embeddings into distance matrices for tree placement.

Protein

SHEST

Samsung Advanced Institute for Health Sciences and Technology / Samsung Medical Center / Sungkyunkwan University

Released November 19, 2025

1

Histopathology model that predicts single-cell type composition and reconstructs spatial gene expression from H&E slides, with no molecular assay.

PathologySpatial omics

H3BERTa

University of Bern

Released November 3, 2025

2151

Antibody language model pretrained only on CDR-H3 loops, giving embeddings for immune repertoire analysis and antibody sequence classification.

ProteinLanguage model

Genos-m

BGI-HangzhouAI

Released May 21, 2026

13026

Mixture-of-Experts genomic foundation model for the human microbiome, with 4.7B parameters pretrained on bacterial, archaeal, and phage genomes.

DNA & Gene

Susagi

University of Zurich

Released May 11, 2026

48

Microbiome world model that treats a community as a set of taxa, scoring how well each member fits and predicting community dynamics zero-shot.

DNA & Gene

PlantCAD2

Cornell University

Released April 3, 2026

4.6K97

Long-context plant DNA language model, 676M parameters on a Mamba2 backbone, pretrained on 65 angiosperm genomes for cross-species variant annotation.

DNA & Gene

Vermeer

Microsoft Research / Broad Institute / Harvard University

Released June 1, 2026

3

Generative microscopy foundation model that synthesizes in-silico fluorescence images of protein subcellular localization from amino-acid sequence.

ImagingProtein

EvoSynth

University of Alabama at Birmingham

Released November 4, 2025

8

Multi-target drug discovery framework pairing a diffusion-transformer generator with evolutionary latent-space search and synthesis-aware scoring.

Small molecule

Mammo-CLIP

Boston University / University of Pittsburgh

Released May 20, 2024

4598

Vision-language foundation model pre-trained on screening mammogram-report pairs to improve data efficiency and robustness in breast cancer detection.

ImagingPathology

BioMatrix

Shanghai AI Laboratory / Renmin University of China

Released June 20, 2026

18441

Decoder-only foundation model that unifies sequences, 3D structures, and natural language for small molecules and proteins in one shared token space.

ProteinSmall moleculeLanguage model

OneGenome-Rice

Zhejiang Lab / BGI Research

Released April 21, 2026

2325

Genomic foundation model for rice, pretrained on 422 Oryza genomes with a 1 Mbp context window and a 1.25B-parameter mixture-of-experts transformer.

DNA & Gene

BioMed Multi-Omic

IBM Research

Released June 1, 2025

2362

Open-source framework for building RNA and DNA foundation models, featuring WCED pretraining for transcriptomics and SNP-aware encoding for genomics.

DNA & Gene

BioT5+

Microsoft Research Asia

Released August 1, 2024

419127

Text-to-text biological language model spanning molecules, proteins, and text, adding IUPAC names and multi-task instruction tuning to BioT5.

Language modelSmall moleculeProtein

BioT5

Renmin University of China

Released October 11, 2023

177127

Encoder-decoder framework unifying molecules, proteins, and natural language with SELFIES notation for cross-modal drug discovery tasks.

Language modelSmall moleculeProtein

scLDM.CD4

Chan Zuckerberg Initiative

Released November 4, 2025

1979

Single-cell latent diffusion model fine-tuned on 14.5 million CD4+ T cells to simulate transcriptomic effects of single-gene perturbations.

Single-cell

ESM Cambrian

EvolutionaryScale

Released December 1, 2024

2.1K2.9K

Protein language model family at 300M, 600M, and 6B parameters, purpose-built for representation learning and outperforming ESM-2 at smaller scale.

Protein

CodonFM

NVIDIA / Arc Institute

Released April 25, 2025

5787

Codon-resolution language models trained on 130 million coding sequences from 20,000 species, learning codon rules for translation and mRNA stability.

RNA

La-Proteina

NVIDIA

Released January 23, 2026

182304

Partially latent flow-matching model for de novo protein design, jointly generating sequence and all-atom structure for proteins up to 800 residues.

Protein

ProstT5

Rostlab

Released July 25, 2023

34.9K318

Bilingual protein language model that translates bidirectionally between amino acid sequences and the 3Di structural alphabet for inverse folding.

Protein

ESM-1b

Meta AI

Released April 5, 2021

4.9K4.2K

Transformer protein language model trained on 250 million protein sequences that learns structural and functional representations without supervision.

Protein

GENA-LM

AIRI Institute

Released June 13, 2023

1.4K230

Family of transformer-based DNA language models using BPE tokenization and BigBird sparse attention to reach context lengths up to 36,000 base pairs.

DNA & Gene