Overview
Zenodo is an open research repository operated by CERN under the OpenAIRE programme, and for a slice of the biological foundation model catalog it is where the trained weights actually live. If you need to download a checkpoint that authors archived alongside their paper — with a permanent DOI, a fixed version, and a citable record — Zenodo is the canonical source. It is not a model runtime or a social hub; it is a durable archive built for reproducibility, which makes it a dependable place to fetch exactly the weights a publication described.
Models with weights on Zenodo
Zenodo hosts downloadable weights across several domains of the catalog. On the protein side this includes EvoDiff, the diffusion-based protein generation model; ProteinBERT, a protein language model; and GearNet and ESM-GearNet, structure-based protein representation models. Nucleic-acid models include RiNALMo, an RNA language model; Orthrus, an mRNA representation model; and TrRosettaRNA for RNA structure. Single-cell and expression models such as SCimilarity and scDiffusion, cancer-imaging biomarker model FMCIB, and the PPG biosignal foundation model PaPaGei are also deposited here. Each resolves to a Zenodo record linked from its bio.rodeo entry; this is a representative selection rather than the complete set.
Downloading weights from Zenodo
Weights are downloaded directly from each record's file listing, either through the web interface or Zenodo's REST API. Because records are versioned and permanently citable, a given DOI always resolves to the same fixed set of files, so a download is reproducible and safe to reference in a methods section. Zenodo runs no hosted inference and no fine-tuning — it is an archive, not a runtime — so the workflow is to fetch the checkpoint and then load, serve, or fine-tune it in your own environment. This suits researchers and agents that value provenance and exact, citation-anchored versioning over a managed serving layer.
Download weights from Zenodo (23)
Protein language model pretrained on UniRef90 with masked language modeling and Gene Ontology annotation prediction, at 16 million parameters.
Geometric relational graph neural network that encodes 3D protein structures through geometry-aware message passing and self-supervised pretraining.
Discrete diffusion model for protein sequence and MSA generation, enabling controllable de novo design directly in sequence space without structure.
RNA 3D structure prediction pipeline pairing a transformer (RNAformer) that predicts inter-nucleotide geometries with Rosetta energy minimization.
FMCIB (Foundation Model for Cancer Imaging Biomarkers)
Harvard Medical School / Dana-Farber Cancer Institute / Brigham and Women's Hospital / Massachusetts General Hospital / Maastricht University / Aarhus University / Stanford University
Released March 15, 2024
Self-supervised 3D CT foundation model that extracts general-purpose tumor representations for cancer imaging biomarker discovery and prognosis.
Deep learning framework predicting equilibrium distributions of molecular systems, enabling efficient ensemble generation and conformation sampling.
RNA language model with 650M parameters pretrained on 36 million non-coding RNA sequences, generalizing structure prediction to unseen RNA families.
Single-cell foundation model trained by metric learning to embed scRNA-seq profiles for cell type annotation and similarity search in cell atlases.
Genomic language model trained on metagenomic scaffolds that learns protein co-regulation and function by modeling gene context and operon structure.
Open foundation model for photoplethysmography (PPG), learning morphology-aware waveform representations for cardiovascular and wearable health tasks.
Diffusion model for synthesizing single-cell RNA-seq data, with guided generation of specific cell types, rare cells, and developmental trajectories.
Joint sequence-structure protein representation framework that fuses ESM-2 language model embeddings with GearNet geometric graph neural networks.
Knowledge-enhanced ECG foundation model aligning a ResNet encoder with LLM-generated disease descriptions for zero- and few-shot interpretation.
TUMSyn
ShanghaiTech University / Hainan University / United Imaging Intelligence
Released September 25, 2024
Text-guided MRI synthesis model that generates brain MR sequences and resolutions on demand from routine scans using imaging-metadata prompts.
Mamba-based mature RNA foundation model, contrastively trained on splice isoforms and 400+ mammalian species orthologs for mRNA property prediction.
Single-lead ECG foundation model pretrained on 12-lead recordings, weighting contrastive pairs by clinical risk for cardiovascular risk prediction.
Transformer model for de novo peptide sequencing that reads amino acid sequences directly from tandem mass spectra, with no protein sequence database.
Structure-conditioned graph transformer trained with masked language modeling to learn residue encodings for inverse folding and antibody design.
Latent diffusion model that designs D-peptide binders against native L-protein targets, generalizing across chirality via axial vector features.
Antibody language model pretrained only on CDR-H3 loops, giving embeddings for immune repertoire analysis and antibody sequence classification.
Protein language model that encodes sequences as discrete words from a learned vocabulary for zero-shot function inference and protein design.
Enhancer RNA mapping model that locates eRNA loci genome-wide from DNA sequence and aggregated RNA-seq signal using a CNN-transformer architecture.
Multi-target drug discovery framework pairing a diffusion-transformer generator with evolutionary latent-space search and synthesis-aware scoring.