All Competitors
Every biological foundation model, evaluated and ranked by the bio.rodeo team
Showing 97–120 of 125 filtered models
ProLLaMA
20795304Protein large language model adapted from LLaMA-2 that unifies sequence generation and superfamily classification in one 7B-parameter framework.
Protein95OpennessscMulan
626—Generative language model for single-cell transcriptomics with 368M parameters, unifying cell type annotation, batch integration, and cell generation.
Single-cell48OpennessRNA-MSM
711091.3KRNA language model trained on multiple sequence alignments of Rfam families, predicting secondary structure and solvent accessibility from homology.
RNA61OpennessProGen2
705——Protein language models from 151M to 6.4B parameters, trained on over a billion sequences for sequence generation and zero-shot fitness prediction.
Protein55OpennessBioT5
127—191Encoder-decoder framework unifying molecules, proteins, and natural language with SELFIES notation for cross-modal drug discovery tasks.
Language modelSmall moleculeProtein74OpennessDARWIN Series
24953—Open large language models for natural science, fine-tuned on physics, chemistry, and materials science literature with automated instruction tuning.
Language model24OpennessTULIP
1343—Unsupervised transformer language model for TCR-epitope binding prediction that generalizes to unseen epitopes without needing negative examples.
Protein60OpennessDNABERT-2
507456170.4KMulti-species genomic foundation model swapping k-mer tokenization for byte pair encoding, matching Nucleotide Transformer with 21x fewer parameters.
DNA & Gene64OpennessLLaVA-Med
2.2K1.9K12.3KBiomedical vision-language assistant for question answering on radiology and pathology images, adapted from LLaVA on PubMed Central captions.
PathologyLanguage model28OpennesstGPT
1762159Tianjin Medical University Cancer Institute and HospitalApril 20, 2023cell_type_annotationfoundation_modellanguage_model+4Single-cell foundation model pre-trained on 22 million transcriptomes, using rank-based gene encoding for clustering and trajectory inference.
Single-cell50OpennessSpecies-Aware DNA LM
29538.1KMasked DNA language model trained on over 800 vertebrate genomes and conditioned on species identity to learn conserved regulatory sequence features.
DNA & Gene76OpennessAnkh
249733.2KParameter-efficient protein language model that matches larger models such as ESM-2 on protein prediction tasks using under 10% of the parameters.
Protein24OpennessSpliceBERT
561—RNA language model pre-trained on 2M+ pre-mRNA sequences from 72 vertebrate species for splice-site prediction and variant effect analysis.
RNA77OpennessReprogBERT
2439—Antibody CDR design model that reprograms a frozen English BERT for sequence infilling, avoiding training a dedicated protein language model.
Protein56OpennessBioGPT
4.5K1.5K101.4KGenerative transformer pretrained on PubMed abstracts for biomedical text generation and mining, including relation extraction and question answering.
Language model66OpennessGenSLM
142142—Genome-scale language model trained on prokaryotic genes and SARS-CoV-2 genomes to model viral evolution and flag emerging variants of concern.
DNA & Gene56OpennessMoLFormer-XL
406595209.9KLarge-scale chemical language model trained on 1.1 billion SMILES strings using linear attention transformers for molecular property prediction.
Small molecule86OpennessESM-2 & ESMFold
4.2K5.1K1.5MMeta AI's family of protein language models scaled to 15B parameters, paired with ESMFold for fast, alignment-free atomic-level structure prediction.
Protein83Openness