InstaDeep / Research Institute of Molecular Pathology (IMP) / Medical University of Vienna / Cornell University / Cold Spring Harbor Laboratory
Released December 22, 2025
Multi-species genomics foundation model spanning representation learning, functional-track prediction, and sequence generation at 1 Mb context.
Bimodal masked language model that jointly encodes bulk RNA-seq expression and DNA methylation into patient-level embeddings for cancer genomics.
Antibody foundation model generating paired sequences with genetic and developability metadata for labelling, humanisation, and library design.
Protein sequence representation model pairing a sequence autoencoder with a denoising diffusion model over its latent space for frozen embeddings.
Transferable coarse-grained force field for molecular dynamics of proteins, RNA, and lipids, built on the MACE equivariant graph architecture.
Bulk RNA-seq foundation model that learns patient-level embeddings from binned gene expression for pan-cancer classification and survival prediction.
Surrogate machine learning force field reusing a reference model's node features from earlier steps, running molecular dynamics eight times faster.
Protein fitness prediction model meta-trained across ProteinGym deep mutational scans, transferring to unseen assays in context with no task data.
Multimodal model fusing pre-trained DNA, RNA, and protein encoders through cross-attention to predict tissue-specific transcript isoform expression.
Conversational agent that answers plain-language questions about DNA, RNA, and protein sequences by coupling a genomics encoder to a language model.
Genome annotation model that labels 14 genic and regulatory element classes at single-nucleotide resolution across DNA windows up to 50 kb.
DNA foundation models from 500M to 2.5B parameters, trained on 3,200+ human genomes and 850 species for variant effect prediction.