A biomedical research institute of MIT and Harvard, uniting genomics, chemistry, and AI to understand the roots of disease and help patients.
Dana-Farber Cancer Institute / Harvard T.H. Chan School of Public Health / Harvard University / Boston Children's Hospital / Massachusetts General Hospital / Harvard Medical School / Broad Institute / Ludwig Center at Harvard
Released August 18, 2026
Pan-cancer clinico-genomic model for treatment response and survival prediction, transferring zero-shot to unseen hospitals and cancer types.
Histopathology foundation model predicting spatial gene expression from H&E slides at single-cell resolution via linear whole-slide attention.
Northeastern University / Broad Institute / KAIST / EPFL / HITS Inc.
Released August 2, 2026
Enzyme-substrate specificity prediction by end-to-end co-folding, with no predefined binding pocket. AUROC 0.766 on unseen enzymes and substrates.
Transformer that predicts protein-protein interactions at residue resolution, spanning mutations, PTMs, peptide-MHC binding, and disease variants.
Generative microscopy foundation model that synthesizes in-silico fluorescence images of protein subcellular localization from amino-acid sequence.
Long-context RNA foundation model that predicts splicing, isoform abundance, and variant effects from 64 kb of unspliced pre-mRNA sequence.
Broad Institute / Howard Hughes Medical Institute / Harvard University / Harvard Medical School / MIT / Massachusetts General Hospital / The Jackson Laboratory / University of Minnesota
Released February 20, 2026
Prime editing efficiency prediction from pegRNA sequence, with every biochemical step of the editing mechanism modeled as its own learned rate.
Technical University of Munich / Helmholtz Munich / Harvard Medical School / Broad Institute / Harvard University
Released October 1, 2025
Predicts single-cell scRNA-seq coverage and scATAC-seq insertion profiles from DNA sequence, adapting the Borzoi trunk with a cell-specific decoder.
Patient-level representation learning from scRNA-seq: a transformer set encoder with a diffusion decoder, fine-tuned for clinical prediction.
Multimodal foundation model predicting genome-wide binding of chromatin-associated proteins from protein sequence, DNA sequence, and chromatin state.
Technical University of Munich / Helmholtz Munich / University of Oxford / Broad Institute
Released July 20, 2025
Splicing variant effect prediction across 49 human tissues and 15 developmental stages, from four weeks post conception to adulthood.
Single-cell perturbation-response model that predicts transcriptomic and imaging outcomes of unseen genetic perturbations via a VAE with attention.
Mahmood Lab / Brigham and Women's Hospital / Harvard Medical School / Broad Institute / Dana-Farber Cancer Institute / Beth Israel Deaconess Medical Center / Stanford University / The Ohio State University / University of Tübingen
Released June 3, 2025
Spatial proteomics foundation model, marker-aware and panel-agnostic, pretrained on 47 million multiplexed tissue-imaging patches from 175 markers.
De novo protein binder design that recasts structure-predictor confidence as an energy function, replacing ipTM as the hallucination objective.
Spatial transcriptomics prediction from H&E whole-slide images. One generative checkpoint covers 38,984 genes and 17 organs without fine-tuning.
Harvard Medical School / AstraZeneca / Broad Institute / Dana-Farber Cancer Institute / Carnegie Mellon University
Released March 4, 2025
Drug-combination safety prediction that fuses molecular structure, pathway knowledge, cell viability, and transcriptomic response to perturbation.
Genomic language model continue-pretrained on 13 million UK Biobank variants, giving variant-aware DNA embeddings for gene function and expression.
Mahmood Lab / Brigham and Women's Hospital / Massachusetts General Hospital / Harvard Medical School / Broad Institute / Dana-Farber Cancer Institute
Released January 28, 2025
Slide-level pathology foundation model that encodes a whole-slide image of any size into one embedding, supervised by paired sequencing data.
University of Washington / Broad Institute / Harvard University / Heidelberg University
Released January 27, 2025
Base-pair resolution sequence-to-activity CNN predicting ATAC-seq Tn5 insertion profiles and accessibility across 90 mouse immune cell types.
University of Toronto / University Health Network / Vector Institute / Broad Institute / Structural Genomics Consortium
Released December 20, 2024
Latent diffusion model that paints high-resolution Cell Painting images of cells responding to a chemical compound or an over-expressed gene.
Mahmood Lab / Mass General Brigham / Harvard Medical School / Brigham and Women's Hospital / Dana-Farber Cancer Institute / Broad Institute / Harvard University / MIT / Helmholtz Munich / Technical University of Munich / Emory University / Pusan National University / University of Tokyo / National Cancer Center Japan
Released December 2, 2024
Histopathology patch encoder turning 512x512 tiles into 768-dimensional features, trained with a CoCa objective on 1.26 million captioned images.
ETH Zurich / SIB Swiss Institute of Bioinformatics / Swiss Data Science Center / EPFL / Dana-Farber Cancer Institute / Broad Institute / Harvard University / Helmholtz Munich / Technical University of Munich
Released November 22, 2024
Enhancer-promoter interaction prediction from DNA sequence and ATAC-seq alone. Spearman above 0.90 on cell types unseen during training.
Conditional autoregressive genomic language model trained on 13.6M mammalian promoters, scoring promoter variants, including indels, zero-shot.
McGill University / Shanghai Jiao Tong University / Mila / Université de Montréal / Hong Kong University of Science and Technology / Institute for Protein Design / Yale University / Northeastern University / Broad Institute / MIT / Google DeepMind
Released November 10, 2024
De novo enzyme design conditioned on the reaction to be catalysed: substrate and product SMILES in, catalytic pocket, enzyme, and docked complex out.
Binding energy for protein-ligand, protein-protein, and antibody-antigen complexes is read off an energy model trained on crystal structures alone.
National Renewable Energy Laboratory / Harvard Medical School / Broad Institute / Dana-Farber Cancer Institute
Released June 22, 2023
Enzyme optimum pH prediction from sequence, ensembling a light-attention network and a support vector regression over frozen ESM-1v embeddings.
Single-cell foundation model pretrained on about 30 million human transcriptomes, using rank-value encoding for context-aware gene network inference.