Zebrafish sequence-to-function model predicting cell-type-specific gene expression from DNA sequence across embryonic development.
Structure-prediction and design engine that turns ESMC sequence representations into all-atom 3D structures of proteins and biomolecular complexes.
Protein language model trained on roughly 2.8 billion sequences, forming the representation core of Biohub's world model of protein biology.
Causal multimodal transformer that embeds the do-operator in attention to predict single-cell gene expression under unseen genetic perturbations.
Masked language model for T-cell receptor and peptide-MHC binding prediction, with compositional pretraining and non-autoregressive decoding.