Every biological foundation model, evaluated and ranked by the bio.rodeo team
Showing 337–360 of 400 filtered models
Genome language model that embeds a genome as a set of contextualized protein embeddings, pretrained on over 100,000 viruses for viromics.
Multi-omics transformer generating transcriptomic, methylation and proteomic signatures for a given tissue, disease, age group, sex and compound.
Chromatin-state language model pretrained on ROADMAP annotations from 127 human cell types to find chromatin-state motifs and predict gene expression.
DNA language model of the human genome with a vocabulary learned by byte-pair encoding rather than fixed k-mers, fine-tuned for genome biology tasks.
Histopathology image translation with diffusion, moving H&E tiles between stains, tumor types, and organ sites and editing them from omics profiles.
RNA language model that switches between nucleotide and byte-pair tokenization by input length, so one 117M encoder handles sequences of any length.
Generative biological foundation model placing DNA, RNA, and protein in one shared vocabulary, spanning genomic, proteomic, and cross-molecule tasks.
Unified DNA, RNA, and protein foundation model with 1.8B parameters, pretrained across 169,861 species to learn the central dogma from sequence.
Interpretable model of human transcription initiation that decomposes promoter activity into a minimal set of sequence rules at base-pair resolution.
Genomic language model trained on metagenomic scaffolds that learns protein co-regulation and function by modeling gene context and operon structure.
Bulk tumor transcriptome model ensembling hundreds of variational autoencoders into interpretable cancer-specific latent spaces for 18 cancers.
Bidirectional, reverse-complement equivariant DNA language models built on Mamba state space models for long-range variant effect prediction.
DNA embedding model built on DNABERT-2, using contrastive learning to cluster sequences by species for metagenomic binning without labeled data.
Nucleotide language model for prokaryotic promoter design, fine-tuned from one pretrained base into 27 species-specific generative checkpoints.
Generative model of bacterial gene content that expands a handful of chosen KEGG modules into the full gene complement a viable cell would need.
DNA language model for variant effect prediction across coding and non-coding regions, using whole-genome alignments of 100 vertebrate species.
Codon-vocabulary protein language model that converts ProtBERT to 64 codon tokens via embedding seeding, masked pretraining, and distillation.
Multi-language transformer framework using five pre-trained language models to predict DNA methylation (6mA, 4mC, 5hmC) across species.
Chromosome-wise explainable autoencoder that compresses DNA methylation array data up to 400-fold while keeping CpG groupings interpretable.