Every biological foundation model, evaluated and ranked by the bio.rodeo team
Showing 313–336 of 400 filtered models
DNA methylation foundation model over 49,156 array CpG sites. Imputes missing values, embeds samples, and predicts epigenetic age and disease risk.
Whole-genome somatic copy-number aberration prediction from bulk RNA-seq alone, with one pan-cancer model covering 33 tumor types.
Poly(A)-tail length change predicted from mRNA 3' UTR sequence in maturing oocytes, scoring how single-nucleotide variants disrupt tail lengthening.
Chromatin accessibility prediction across 29 human immune cell types, using in-silico saturated mutagenesis to score constrained regulatory regions.
Frugal conditional diffusion model that samples human SNP haplotypes in PCA space, producing artificial genomes for a chosen continental ancestry.
DNA methylation foundation model reconstructing genome-wide profiles from sparse input. Outperforms GrimAge2 aging clocks on mortality prediction.
Genomic language model that labels adapter sequences in nanopore direct-RNA reads base by base, then splits the chimeric reads those adapters create.
Genomic sequence classification answered through natural-language prompts by one GPT-2 pretrained on mixed DNA and English under one BPE vocabulary.
Autoregressive temporal convolutional network for synthetic yeast promoter design, trained with guidance from a sequence-to-expression predictor.
Context-only BERT for bacterial protein function prediction, reading genomes as sentences of protein-cluster tokens with no sequence input.
Mixed-modal DNA, RNA, and protein foundation model at 110M and 270M parameters, with in-context learning across sequence modalities.
Joint embedding space for common SNPs and free-text clinical concepts, aligned by contrastive learning over GWAS, biobank and knowledge-graph pairs.
Codon-resolution language model suite pairing a bidirectional encoder with an autoregressive decoder over protein-coding sequences.
Perturbation target identification for single-cell transcriptomics, reading intervened genes off the difference between two inferred causal graphs.
Cis-regulatory element classifier that reads DNA sequence plus chromatin accessibility and loop tracks to label enhancers, silencers and insulators.
RNA language model that predicts G-quadruplex formation and subtype from transcript sequence and scores how single-nucleotide variants alter folding.
Gene expression prediction across a megabase of DNA that aligns frozen regulatory sequence features to language-model tokens by cross-attention.
Generative DNA language model for plasmid design and annotation, pretrained on 153,208 engineered plasmid sequences deposited in Addgene.
Manufacturing-aware generative sequence models whose parameters are DNA synthesis reaction conditions, so designs are made in vitro at petascale.
Nanopore basecaller for fully 5-hydroxymethylcytosine-substituted DNA, reading raw ion current from strands that standard basecallers cannot resolve.
Conditional generator of bulk transcriptome and DNA methylation profiles, sampling tissue-, age- and species-matched synthetic omics samples.
A Wasserstein GAN that generates artificial human genomes in PCA space, synthesizing 65,535-SNP haplotypes for 26 worldwide populations.