Every biological foundation model, evaluated and ranked by the bio.rodeo team
Showing 193–216 of 400 filtered models
Distribution-level representation learning that embeds whole cell populations, perturbation responses, and sequence sets, not individual data points.
Bidirectional DNA foundation model with a Mamba-attention-mixture-of-experts design, reading 1 million base pairs at single-nucleotide resolution.
Contrastive k-mer embedding model for sequencing reads whose latent space encodes genomic position, matching BWA-aln accuracy on ancient DNA mapping.
Transcriptomic perturbation prediction across unseen single and double gene knockdowns and unseen cell lines, driven by gene-gene knowledge graphs.
Protein-DNA binding prediction and binder design from sequence, aligning protein and DNA language model embeddings instead of co-folding a complex.
Tri-modal pathology foundation model aligning whole-slide images, transcriptomes, and diagnostic reports, and running on any subset of the three.
Single-cell chromatin accessibility foundation model with genome-aware tokenization, pretrained on 1.97 million scATAC-seq cells across 30 tissues.
Genomic prediction model for plant and animal breeding, pretrained entirely on simulated populations and deployed with no training or tuning.
Single-cell foundation model inferring cis-regulatory relationships from scRNA-seq and scATAC-seq, pretrained on an atlas of 1.3 million cells.
Alignment-free biosynthetic gene cluster detection and annotation from ESM-2 gene embeddings in genomic context, up to 102x faster than antiSMASH.
CRISPR/Cas9 off-target prediction that fine-tunes a DNA language model and gates in chromatin signal, reaching 0.550 PR-AUC on GUIDE-seq data.
Whole-genome bacterial pathogenicity prediction from ProtT5 embeddings, alignment-free and taxonomy-agnostic, with per-protein attention scores.
Multi-modal single-cell foundation model that projects Enformer DNA embeddings into a transcriptome model token space to predict gene regulation.
Multimodal 3D genome foundation model pairing Hi-C contact maps with epigenomic tracks, pretrained on over one million paired samples.
C/D box snoRNA gene predictor for any eukaryote genome, built on DNABERT and able to separate expressed snoRNAs from their pseudogenes.
Bacterial genome language model tokenizing whole genomes as ordered conserved elements; frozen embeddings beat Pfam baselines on 23 of 25 phenotypes.
DNA and RNA language model with a data-driven 4,096-token unigram vocabulary, matching larger genomic foundation models at 89.2M parameters.
Regulatory genomics foundation model pretrained on 6,391 human ChIP-seq cistromes, representing how ~1,000 transcription regulators cooperate.
Hi-C resolution enhancement model combining a U²-Net with self-attention to recover TAD boundaries and chromatin loops from sparse contact maps.
Enhancer models predicting cell-type-specific chromatin accessibility from DNA sequence, with a pretrained zoo and synthetic enhancer design tools.
Protein function prediction fusing five Gene Ontology pipelines, two of them deep models over protein and DNA language model embeddings.
DNA language model for SELEX aptamer libraries that embeds single-stranded oligonucleotides so enrichment and target specificity become measurable.