Protein language model family from 8M to 15B parameters, used as a frozen sequence encoder whose representations encode atomic-level structure.
Chromatin-level variant effect prediction from DNA sequence, projecting 21,907 predicted regulatory profiles onto 40 interpretable sequence classes.
Renal pathology segmentation covering six kidney tissue types across 5x to 40x magnifications from one network, and transferring from human to mouse.
RNA methylation site predictor combining multiple sequence encodings with dilated convolution and BiLSTM layers to identify m6A and m1A sites.
Glomerular detection, segmentation, and glomerulosclerosis subtyping from renal whole-slide images, pretrained on web-mined glomerular figures.
Autoregressive protein language model scoring variant effects zero-shot, blending sequence likelihood with homolog statistics retrieved at inference.
Protein language model family built on CNNs rather than transformers, matching transformer quality while scaling linearly with sequence length.
Predicts 3D genome architecture directly from DNA sequence across nine scales, from 4-kb contacts up to a 256-Mb whole-chromosome window.
Gut microbiome taxa embeddings that project a 16S V4 ASV table into a shared property space so classifiers transfer between cohorts.
Full-atom protein model accuracy estimation, regressing per-atom lDDT with an SE(3)-transformer over a heavy-atom graph of the modeled structure.
Resolution enhancement for Hi-C contact matrices, reconstructing full-depth 10 kb maps from libraries sequenced at a fraction of the read depth.
BERT-based predictor of DNA N6-methyladenine (6mA) modification sites, using word2vec encoding and cross-species transfer learning.
RNA language model that learns base-level embeddings capturing sequence context and secondary structure, enabling fast structural alignment.
SMILES language model pretrained on 100M molecules, transferring to forward reaction prediction, retrosynthesis, molecular optimisation, and QSAR.
Protein language model that fuses Gene Ontology knowledge graphs with masked language modeling, improving protein function and interaction prediction.
Protein language model pretrained on UniRef90 with masked language modeling and Gene Ontology annotation prediction, at 16 million parameters.
Transformer that imputes missing CpG methylation states from sparse single-cell bisulfite sequencing, modeling genomic and cell-level structure.
Antibody-specific language model trained on the OAS database for restoring missing residues and generating high-quality sequence representations.
Medical-domain CLIP fine-tuned on radiology image-caption pairs from ROCO, serving as a drop-in visual encoder for medical visual question answering.