Every biological foundation model, evaluated and ranked by the bio.rodeo team
Showing 169–192 of 400 filtered models
Species-conditioned codon language model that jointly reads 5'UTRs, coding sequence, and RNA secondary structure to design native-like genes.
Controllable DNA sequence design conditioned on cell type, transcription factor, or activity signal, in GPT- and BERT-style transformer variants.
Bacterial promoter annotation and expression prediction from a 1.8M-parameter transformer pretrained on 9M gammaproteobacterial regulatory regions.
Protein-DNA binding free energy change prediction for missense mutations, with double- and single-stranded DNA binders modeled separately.
Fungal genome mining framework that detects biosynthetic gene clusters and identifies their core enzymes from a pretrained Pfam-domain transformer.
Tissue-specific RNA splicing prediction from pre-mRNA sequence, scoring how variants shift splice-site usage across 18 human tissues.
All-atom biomolecular structure prediction with adapters for allosteric states, user-defined interface constraints, and binding affinity.
DNA foundation model that predicts thousands of functional genomic tracks, from expression and splicing to chromatin, at single base-pair resolution.
Plasmid characterization and retrieval model aligning DNA sequences with property text across ten facets, from antimicrobial resistance to host range.
Generative codon language model for mRNA design, trained on 338,417 coding sequences with inference-time masking that preserves the encoded protein.
Decoder-only transformer that recasts ancestral recombination graph inference as next-token prediction, estimating coalescence times from variation.
Alignment-free taxonomic classification of eukaryotic DNA in metagenomes. Reaches an F1 of 0.871 on 500 bp contigs, where k-mer tools falter.
Genome-anchored histopathology embeddings that predict molecular biomarkers, subtypes, and survival from whole-slide images alone at inference.
Bulk transcriptome foundation model, 150M parameters over ~20,000 protein-coding genes. Imputes masked expression at Pearson r = 0.954.
Multimodal viral foundation model over nucleotide and protein sequence, built for virus discovery, function annotation, and antibody design.
Multimodal large language model that writes free-form natural-language gene function descriptions directly from a nucleotide sequence and a prompt.
Domain-level language model treating Pfam protein domains as tokens to predict and design bacterial and fungal biosynthetic gene clusters.
Codon optimization model for heterologous expression in E. coli, fine-tuning ProtBert to label each residue with an expression-weighted codon.
DNA foundation model pretrained by supervised genomic profile prediction, using mixture-of-experts routing across species and assay types.
Open-source framework for building RNA and DNA foundation models, featuring WCED pretraining for transcriptomics and SNP-aware encoding for genomics.
Cross-modal co-embedding of biosynthetic gene clusters and natural products, enabling bidirectional retrieval between gene cluster and compound.
DNA-LLM reasoning model fusing genome foundation model embeddings with an LLM to produce step-by-step pathway and variant effect explanations.
Compact 1.1M-parameter DNA language model distilled from Nucleotide Transformer v2, outperforming its 500M teacher on 11 of 18 benchmark tasks.