Every biological foundation model, evaluated and ranked by the bio.rodeo team
Splice-site and variant-impact prediction from DNA sequence, with pretrained models for human, mouse, zebrafish, honey bee, and Arabidopsis.
Long-context genomic foundation model reading up to 192,000 base pairs at single-nucleotide resolution with dense LLaMA-style self-attention.
Bacterial genome encoder that renders draft assemblies as Chaos Game Representation images and embeds them for nearest-neighbour species search.
Enhancer prediction from DNA sequence alone, reaching 88.05% accuracy and 76.22% MCC on held-out ENCODE cCRE regions and annotating the whole genome.
Natural product chemistry prediction from biosynthetic gene clusters, assigning ChemOnt ontology classes from the cluster's Pfam domain composition.
Phylogenetic tree inference from unaligned nucleotide sequences, using a 2D genomic-footprint encoding and CNN classification of triplet topologies.
Genomic language models with disentangled attention, pretrained on prokaryotic and eukaryotic genomes for sequence classification and variant effects.
Genomic language model for variant effect prediction that scores deleterious mutations from a single DNA sequence, with no alignment at inference.
Secondary metabolite structure prediction from microbial biosynthetic gene clusters, generating SMILES strings from Pfam functional-domain tokens.
Structure-aware adapter that injects chromatin-derived gene regulatory networks into single-cell RNA foundation models like scGPT and scFoundation.
Genomic language model continue-pretrained on 13 million UK Biobank variants, giving variant-aware DNA embeddings for gene function and expression.
Multimodal sequence model spanning proteins, coding DNA, and regulatory DNA for zero-shot fitness scoring and conditional sequence generation.
Prophage island detection in bacterial genomes and metagenome-assembled genomes, pairing a fine-tuned ESM-2 gene classifier with density clustering.
Histopathology model predicting gene expression and DNA methylation from H&E slides across 23 cancer types, fusing FFPE and fresh-frozen predictors.
Predicts haplotype-specific 3D genome organization and Hi-C contact maps from a single long-read Fiber-seq assay, using no DNA sequence as input.
DNA barcode foundation model whose masked-autoencoder pretraining keeps mask tokens out of the encoder, for arthropod taxonomic classification.
Chromatin signal prediction that fuses reference DNA with per-nucleotide ATAC-seq, calling CTCF and histone marks in any cellular context.
Single-cell DNA methylation foundation model capturing genome-wide CpG dependencies in whole-genome bisulfite sequencing across tissues and species.
Sequence-to-chromatin models for cattle, chicken, pig, and Atlantic salmon that score the regulatory impact of non-coding variants genome-wide.
Genomic foundation model trained on 9.3 trillion DNA base pairs across all domains of life, with 40B parameters and a 1-million-token context.
Functional genomics models predicting RNA-seq, CAGE, and DNase coverage tracks from DNA sequence using striped bidirectional Mamba and attention.
Long-range DNA language model interleaving attention with Mamba2 state-space layers to read 131kb of sequence at single-nucleotide resolution.
Cell-free DNA methylation deconvolution at individual-read resolution, estimating cell-type proportions and condition-specific methylation profiles.