A generative biology company building multiscale foundation models that span molecules, cells, and phenotypes toward an AI-driven digital organism.
Sparse mixture-of-experts DNA encoder pretrained on multi-species genomes at 235M to 1B parameters, reading 64 kb of context and 256 kb at 235M.
Transformer U-Net pretrained on 6 trillion tokens of multi-species DNA, predicting expression and epigenomic tracks across 1 Mb of context.
World model that simulates a human cell as one persistent state, propagating drug and gene perturbations from DNA through to whole-cell morphology.
Histopathology foundation model with 1.1B parameters, trained entirely on public data using JEDI, a dual-stage strategy combining JEPA and DINO.
The Hong Kong Polytechnic University / genbio.ai / Mohamed bin Zayed University of Artificial Intelligence
Released December 17, 2025
Multimodal protein language model that adds a continuous-token diffusion head to a discrete pLM, modeling structure without vector quantization.
genbio.ai / Mohamed bin Zayed University of Artificial Intelligence / Carnegie Mellon University
Released July 7, 2025
Spatial transcriptomics foundation model pretrained on 22 million cells, encoding each cell with its neighbors for niche and density prediction.
genbio.ai / Mohamed bin Zayed University of Artificial Intelligence / Carnegie Mellon University
Released December 5, 2024
Protein structure tokenizer that discretizes backbones into 512 discrete tokens and reconstructs all-atom structures, including side chains.
DNA foundation model scaling an encoder-only transformer to 7 billion parameters for variant effect prediction, gene expression, and sequence design.
Mixture-of-experts protein language model scaling to 16 billion parameters, applied to variant effect prediction and de novo protein design.
Single-cell RNA-seq foundation model pretrained on 50 million human cells, encoding the full transcriptome for annotation and perturbation modeling.
RNA foundation model with 1.6 billion parameters, pretrained on 42 million non-coding RNA sequences for structure prediction and RNA sequence design.