An independent nonprofit research institute in Palo Alto pairing curiosity-driven biology with AI to understand and treat complex human disease.
Generative pipeline for epitope-targeted de novo antibody (nanobody) CDR design that yields nanomolar binders from only dozens of designs per antigen.
Multimodal reasoning LLM for protein function prediction, fusing protein language model embeddings to emit interpretable GO-term reasoning traces.
Tokenizer-free genomic foundation model that adaptively chunks raw nucleotides, enabling zero-shot variant fitness and gene essentiality prediction.
Single-cell foundation model using tabular attention over context cells to predict responses to arbitrary perturbations without fine-tuning.
Virtual cell transformer that predicts how cells respond to genetic, chemical, or signaling perturbations, generalizing to unseen cellular contexts.
Bowang Lab / University of Toronto / Vector Institute / University Health Network / Arc Institute / UCSF
Released May 29, 2025
DNA-LLM reasoning model fusing genome foundation model embeddings with an LLM to produce step-by-step pathway and variant effect explanations.
Codon-resolution language models trained on 130 million coding sequences from 20,000 species, learning codon rules for translation and mRNA stability.
Cell-free RNA language model for liquid biopsy, fusing RNA sequence embeddings with cfRNA abundance across 306 billion pretraining tokens.
Genomic foundation model trained on 9.3 trillion DNA base pairs across all domains of life, with 40B parameters and a 1-million-token context.
Fixed-backbone protein sequence design that co-generates amino acid identity and sidechain conformation, with 49.7% sequence recovery on CATH 4.2.
Bowang Lab / University Health Network / University of Toronto / Vector Institute / Arc Institute / UCSF
Released February 8, 2025
Spatial transcriptomics foundation model continually pretrained on 30 million profiles, with a protocol-aware mixture-of-experts decoder.
Genomic foundation model with 7B parameters that models prokaryotic DNA, RNA, and protein at single-nucleotide resolution over a 131k-token context.
Codon-resolution language model suite pairing a bidirectional encoder with an autoregressive decoder over protein-coding sequences.
Structure-conditioned protein language model aligned to experimental stability data, scoring variant stability and generating stabilized sequences.