bio.rodeo
ModelsOrganizationsLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
RNA foundation models
RNA

RfamGen

Kyoto University / Waseda University

Generative RNA design model that samples family sequences from a VAE latent space constrained by Rfam covariance models and consensus structure.

Released: January 2024

RfamGen is a deep generative model for the data-efficient design of functional RNA family sequences. Developed by researchers at Kyoto University and Waseda University, the model addresses a core challenge in RNA design: most generative approaches treat sequences as simple strings and discard the structural and evolutionary context encoded in multiple sequence alignments (MSAs). RfamGen instead builds those constraints directly into its architecture by grounding the generative process in covariance models (CMs), probabilistic representations that jointly capture sequence conservation and consensus secondary structure across RNA families.

The model frames RNA design as a learned sampling problem over a continuous latent space. By encoding alignment features derived from CMs into a Variational Autoencoder (VAE), RfamGen learns a semantically structured representation in which nearby points in latent space correspond to sequences with similar functional and structural properties. Novel sequences are generated by sampling from this latent space and decoding through the CM, which constrains outputs to respect the base-pairing and conservation patterns that define a given RNA family.

RfamGen was validated across 18 diverse RNA families drawn from the Rfam database, each with alignments of at least 10,000 sequences. In a key experimental test, RfamGen-designed ribozyme sequences demonstrated measurable enzymatic activity as assayed by quantitative massively parallel assays, while randomly sampled sequences from the same families did not. This direct functional validation distinguishes RfamGen from models evaluated solely on sequence statistics.

#Key Features

  • Covariance model VAE: Integrates CM-based encoding directly into the VAE architecture, embedding both sequence conservation and RNA secondary structure constraints into the generative process rather than treating them as post-hoc filters.
  • Structure-aware latent space: The learned latent representation captures functionally relevant sequence variation, enabling smooth interpolation and targeted sampling toward sequences with desired properties.
  • Data-efficient design: Explicit use of MSA and secondary structure information allows the model to learn from fewer examples than purely sequence-based generative models, making it viable for RNA families with limited characterized members.
  • Experimentally validated functionality: Ribozyme sequences generated by RfamGen exhibit catalytic activity measured through quantitative massively parallel assays, providing direct wet-lab evidence of biological utility.
  • Broad RNA family coverage: Demonstrated on 18 distinct RNA families spanning diverse structural classes, including ribozymes, riboswitches, and structural RNAs.

#Technical Details

RfamGen employs a VAE framework in which both the encoder and decoder are built around covariance models. A CM is a probabilistic graphical model on a tree structure that combines a profile hidden Markov model with RNA secondary structure, representing paired and unpaired positions in an RNA consensus fold. This representation allows RfamGen to formalize an MSA under the constraint of a known secondary structure, capturing co-evolutionary signals between paired bases that are essential for functional RNA design.

During training, the model vectorizes alignment features derived from a target RNA family's CM and learns to map these features to a continuous Gaussian latent space using the VAE objective. Sampling from the prior and decoding through the CM-based decoder generates novel paths on the model — equivalent to sampling aligned sequences that respect the structural grammar of the RNA family. Training data is drawn entirely from Rfam, a curated database of non-coding RNA families that provides MSAs, consensus secondary structures, and covariance models.

Evaluation against baseline generative approaches across 18 RNA families showed that RfamGen consistently produced sequences that more closely recapitulate the statistical properties of natural family members. The experimental ribozyme design experiment provided functional confirmation that the latent space encodes biologically meaningful variation beyond sequence-level statistics.

#Applications

RfamGen is primarily applicable to RNA synthetic biology and engineering tasks where functional sequence design is the goal. Researchers working on ribozyme engineering can use the model to generate candidate catalytic RNA sequences with predictable structural characteristics, reducing the combinatorial search space before experimental screening. The model is relevant to RNA-based therapeutics development, where novel non-coding RNA sequences with specific structural scaffolds are needed. More broadly, RfamGen provides a framework for exploring the sequence space of any Rfam-catalogued RNA family in a structurally guided way, supporting evolutionary studies of sequence-structure-function relationships and the development of RNA aptamers or regulatory elements for synthetic genetic circuits.

#Impact

RfamGen was published in Nature Methods in 2024 and represents one of the first generative models for RNA design to achieve direct experimental functional validation at scale. By demonstrating that a structure-aware VAE can produce ribozymes with genuine catalytic activity, the work establishes a benchmark for what constitutes meaningful success in computational RNA design, shifting evaluation from sequence statistics toward functional assays. The covariance model integration is a methodologically distinct contribution that may influence future generative approaches to structured RNA families. A notable limitation is that RfamGen requires a well-characterized Rfam family with a high-quality CM and a sufficient number of aligned sequences, which constrains its immediate applicability to RNA families that lack deep Rfam coverage or for which no consensus secondary structure exists.

Citation

Deep generative design of RNA family sequences

Sumi, S., Hamada, M. & Saito, H. Deep generative design of RNA family sequences. Nat Methods 21, 435–443 (2024).

DOI: 10.1038/s41592-023-02148-8

Recent citations

Papers that recently cited this model.

  • BindRNAgen: Protein-binding RNA sequence generation using latent diffusion models.

    Yan Zhou, Xiaojian Liu, Shengfan Wang, et al.

    Journal of Molecular Biology · Jul 2026

    0
  • Unlocking Your Programmable and Creative RNA Sequence Designer with RDiffusion

    Jue Wang, Jintong Dong, Tianhao Li, et al.

    bioRxiv · Jun 2026

    0Influential
  • The Use of Deep Learning in RNA Therapeutic Development.

    D. Subramanian, Sophia L Yao, Alvin Chan, et al.

    ACS Nano · Jun 2026

    0

Top citations

The most-cited papers that cite this model.

  • Foundation models in bioinformatics

    Fei Guo, Renchu Guan, Yaohang Li, et al.

    National Science Review · Jan 2025

    44
  • GenerRNA: A generative pre-trained language model for de novo RNA design

    Yichong Zhao, Kenta Oono, Hiroki Takizawa, et al.

    bioRxiv · Feb 2024

    41
  • RNA function follows form – why is it so hard to predict?

    Diana Kwon

    Nature · Mar 2025

    22
  • Deep learning for RNA structure prediction.

    Jiuming Wang, Yimin Fan, Liang Hong, et al.

    Current Opinion in Structural Biology · Feb 2025

    19
  • Engineering circular RNA medicines

    Xiaofei Cao, Zhengyi Cai, Jinyang Zhang, et al.

    Nature Reviews Bioengineering · Nov 2024

    19

Related models

Models with similar goals, methods, or subject matter.

  • GenerRNA

    Preferred Networks

    Transformer-based generative language model for de novo RNA design, pretrained on 16 million non-coding RNA sequences from RNAcentral.

    RNA
  • yakRNA Design

    Stanford University

    110M-parameter RNA language model that designs sequences from secondary structure, motif, and Gene Ontology constraints via discrete diffusion.

    RNA
  • RDiffusion

    Zhejiang University / Nanjing Tech University / Sichuan Agricultural University / Shanghai Institute of Materia Medica / Nanjing University / Chinese Academy of Sciences / University of Chinese Academy of Sciences

    Diffusion-based generative RNA model for de novo sequence design, conditioned on function, RNA family, structure, or binding proteins.

    RNA
  • gRNAde

    MRC Laboratory of Molecular Biology / University of Cambridge

    RNA inverse-folding model that generates sequences predicted to fold into a target 3D backbone, capturing non-canonical pairs and tertiary motifs.

    RNA
  • EVA

    GENTEL Lab

    Generative RNA foundation model trained on 114 million full-length sequences for de novo design of tRNAs, aptamers, CRISPR guide RNAs, and mRNAs.

    RNA

Citations

Total Citations60
Influential5
References64

GitHub

Stars42
Forks12
Open Issues2
Contributors1
Last Push3mo ago
LanguageJupyter Notebook

Fields of citing research

  • Biology88%
  • Computer Science85%
  • Medicine68%
  • Engineering13%
  • Chemistry7%
  • Environmental Science3%
  • Materials Science3%
  • Physics3%

Share of papers citing this model.

Openness

bio.rodeo opennessClosed · low usability and reproducibility
10Closed
Usability — can I run it?10
Reproducibility — can I retrain it?10
Model Openness Framework
Unclassified
Restrictive license on core components

Tags

foundation_modelsequence_designvariational_autoencoder

Resources

GitHub RepositoryResearch PaperDataset