An Ivy League research university in New Haven, Connecticut, spanning undergraduate, graduate, and professional education from medicine to the arts.
Genomic foundation model that pairs 650 kb of gene-centered DNA with transcription factor activity to predict expression in unseen cell types.
Contrastive multimodal model for perturbation screens, aligning transcriptomic signatures with text and cell-painting image embeddings.
Duke-NUS Medical School / Genome Institute of Singapore / National University of Singapore / National Cancer Centre Singapore / Yale University
Released September 19, 2025
Single-cell foundation model domain-adapting Llama-3.1-8B on 1.3M gastric cancer cells with gene-family cell sentences instead of ranked-gene order.
Multimodal transformer predicting alternative splicing outcomes across C. elegans neuron subtypes, reaching Spearman ρ = 0.88 on held-out exons.
Paige AI / Microsoft Research / Memorial Sloan Kettering Cancer Center / Yale University
Released June 16, 2025
Multimodal slide-level pathology foundation model trained by clinical-dialogue supervision on 2.3M whole-slide images and 14M Q&A pairs.
Spatial omics foundation model that represents tissue as a hierarchical graph of neighboring cells over per-cell gene co-expression networks.
Spatial transcriptomics prediction from H&E whole-slide images. One generative checkpoint covers 38,984 genes and 17 organs without fine-tuning.
Single-cell foundation model reading scRNA-seq profiles as ranked gene-name sentences, scaled on Gemma-2 for annotation, reasoning and drug screens.
Genomic language model continue-pretrained on 13 million UK Biobank variants, giving variant-aware DNA embeddings for gene function and expression.
Protein-ligand binding affinity prediction from multimodal representations. Retains accuracy on predicted rather than crystal complex structures.
Single-cell RNA-seq representation model that separates batch-dependent from batch-independent variation to compare disease states across datasets.
Spatial gene expression prediction from H&E tumor histology, aligning a pathology foundation model with a single-cell RNA-seq foundation model.
Patient-level foundation model that pools every cell in an scRNA-seq sample into one disease representation, trained on 24.3 million cells.
McGill University / Shanghai Jiao Tong University / Mila / Université de Montréal / Hong Kong University of Science and Technology / Institute for Protein Design / Yale University / Northeastern University / Broad Institute / MIT / Google DeepMind
Released November 10, 2024
De novo enzyme design conditioned on the reaction to be catalysed: substrate and product SMILES in, catalytic pocket, enzyme, and docked complex out.
Biochemistry-aware inverse folding model that augments backbone geometry with physicochemical point clouds, reaching ~90% sequence recovery on CATH.
Peptide-MHC class I immunogenicity prediction fusing sequence, predicted structure, and biochemical properties for vaccine and neoantigen design.
Ataraxis AI / NYU Grossman School of Medicine / New York University / Karmanos Cancer Institute / Cancer Center Baselland / The Catholic University of Korea / University of Chicago / Memorial Sloan Kettering Cancer Center / University of Aberdeen / University Hospital Basel / Cancer Research Malaysia / Omica.bio / Vilnius University / Providence / UPMC Hillman Cancer Center / Northwell Health / Yale University / Meta AI
Released October 28, 2024
Breast cancer recurrence risk read from an H&E slide and six routine clinical variables, using a frozen pan-cancer pathology encoder.
Framework turning single-cell expression profiles into ranked gene-name sequences, letting off-the-shelf language models generate and annotate cells.
University of Washington / Institute for Protein Design / Fred Hutchinson Cancer Center / Yale University / MIT
Released July 9, 2024
Structure-based mutational effect prediction from local atomic environments, scoring how substitutions change protein stability and binding affinity.
Drug combination synergy prediction across cancer cell lines, from LLM text embeddings of drugs and cell lines rather than structures or expression.
Bowang Lab / University Health Network / University of Toronto / Vector Institute / Western University / New York University / Yale University
Released January 22, 2024
Promptable foundation model for universal medical image segmentation, fine-tuned from SAM on 1.57M image-mask pairs across 10 imaging modalities.
Yale University / Baylor College of Medicine / Princeton University
Released September 12, 2023
fMRI foundation model pretrained with masked autoencoding on roughly 6,700 hours of recordings for clinical prediction and network discovery.