bio.rodeo
ModelsOrganizationsLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Single-cell foundation models
Single-cell

scMulan

Tsinghua University

Generative language model for single-cell transcriptomics with 368M parameters, unifying cell type annotation, batch integration, and cell generation.

Released: January 2024
Parameters: 368 Million

scMulan is a multitask generative pre-trained language model developed at Tsinghua University for comprehensive single-cell transcriptomic analysis. Released in early 2024, the model addresses a fundamental challenge in the field: most existing single-cell tools are designed for individual tasks, requiring researchers to assemble and coordinate multiple specialized models to complete a typical analysis workflow. scMulan replaces this fragmented approach with a single, unified framework capable of performing cell type annotation, batch integration, and conditional cell generation within the same model.

The core innovation is a structured cell representation scheme called "cell sentences" (c-sentences). Rather than treating a cell's transcriptome as a flat vector of gene expression values, scMulan encodes each cell as a sequence of tuples, where each element pairs an entity (a gene, a metadata term, or a task specification) with its corresponding value. This representation allows the model to incorporate biological metadata — such as tissue of origin, experimental condition, and assay type — directly into the input, giving it the contextual grounding needed to generalize across diverse datasets and experimental settings.

Trained on 10 million single-cell transcriptomic profiles with associated metadata, scMulan's 368 million parameters capture both fine-grained gene regulatory relationships and higher-level tissue-level patterns. Task-specific behavior is controlled through natural-language-style task prompts, enabling zero-shot inference without any additional fine-tuning.

#Key Features

  • Unified Multitask Framework: A single model handles cell type annotation, batch integration, and conditional cell generation, eliminating the need to coordinate separate specialized tools.
  • Zero-Shot Task Execution: Tasks are directed via structured prompts at inference time, so the model generalizes to new datasets and conditions without retraining or fine-tuning.
  • C-Sentence Cell Representation: Cells are encoded as ordered sequences of entity-value tuples covering gene expression, metadata terms, and task descriptors, enabling rich contextual learning.
  • Multi-Organ Annotation Coverage: Pre-configured for zero-shot cell type annotation across seven major human organs: heart, lung, liver, bone marrow, blood, brain, and thymus.
  • Metadata-Aware Pre-Training: Biological metadata is integrated during pre-training rather than treated as auxiliary information, making the model robust to batch effects and study-level variation.

#Technical Details

scMulan is a generative transformer language model with 368 million parameters, pre-trained on 10 million human single-cell RNA-seq profiles sourced from publicly available atlases spanning multiple tissues and experimental protocols. Input cells are formatted as c-sentences: structured sequences of (entity, value) tuples encoding gene expression levels, associated metadata fields (tissue, condition, donor), and the desired downstream task. This unified tokenization scheme allows the same forward pass to serve annotation, generation, and integration objectives, with task selection controlled by a prompt prepended to each c-sentence.

Pre-training employed a multi-task objective designed to simultaneously learn gene expression patterns, metadata associations, and generative dynamics. The result is a model that can produce coherent cell-state representations in a shared embedding space suitable for downstream tasks including UMAP visualization, nearest-neighbor classification, and batch-corrected integration. For zero-shot annotation across the seven supported organs, the model leverages task prompts that specify the target tissue and annotation granularity, then generates cell type labels autoregressively from the learned distribution.

#Applications

scMulan is intended for researchers working with human single-cell RNA-seq data who need to annotate cell types, integrate datasets from different experimental batches, or generate synthetic transcriptomic profiles for data augmentation or in silico perturbation studies. Its zero-shot annotation capability is particularly useful for atlas-scale projects where manually curated reference labels are unavailable or incomplete. The batch integration function allows harmonization of data across studies with differing library preparation protocols, sequencing platforms, or donor backgrounds. Conditional cell generation opens avenues for simulating cellular responses to perturbations or for balancing underrepresented cell populations in training datasets for downstream classifiers.

#Impact

scMulan represents an important step toward truly general-purpose foundation models for single-cell biology, demonstrating that a single generative model can unify tasks that have historically required separate tools. By framing single-cell analysis as a language modeling problem over structured cell sentences, the work establishes a principled architecture for incorporating heterogeneous metadata into transcriptomic models. At the time of release, the approach was novel in its simultaneous support for discriminative (annotation), integrative (batch correction), and generative tasks within one pre-trained checkpoint. The model's current scope is limited to seven human organs and does not yet cover non-human species or non-transcriptomic modalities such as chromatin accessibility or protein expression, areas that represent natural extensions for future work.

Citation

scMulan: a multitask generative pre-trained language model for single-cell analysis

Preprint

Bian, H., Chen, Y., Dong, X., Li, C., Hao, M., Chen, S., Hu, J., Sun, M., Wei, L., & Zhang, X. (2024). scMulan: a multitask generative pre-trained language model for single-cell analysis. bioRxiv, 2024.01.25.577152.

DOI: 10.1101/2024.01.25.577152

Recent citations

Papers that recently cited this model.

  • Benchmarking single-cell foundation models for real-world RNA-seq data integration

    Siyu Han, Tamas Sztanka-Toth, Enes Senel, et al.

    bioRxiv · Apr 2026

    0
  • Representation learning of single-cell RNA-seq data

    Constantin Ahlmann-Eltze, Florian Barkmann, Jan Lause, et al.

    RNA: A publication of the RNA Society · Jan 2026

    1
  • hECA v2.0: an AI-ready ensemble cell atlas of single-cell RNA and ATAC sequencing data

    Xi Xi, Yixin Chen, Xinze Wu, et al.

    Scientific Data · Dec 2025

    1

Top citations

The most-cited papers that cite this model.

  • Large Language Models in Drug Discovery and Development: From Disease Mechanisms to Clinical Trials

    Yizhen Zheng, Huan Yee Koh, Maddie Yang, et al.

    arXiv.org · Sep 2024

    57
  • scGenePT: Is language all you need for modeling single-cell perturbations?

    Ana-Maria Istrate, Donghui Li, Theofanis Karaletsos

    bioRxiv · Oct 2024

    12
  • Representation learning of single-cell RNA-seq data

    Constantin Ahlmann-Eltze, Florian Barkmann, Jan Lause, et al.

    RNA: A publication of the RNA Society · Jan 2026

    1
  • hECA v2.0: an AI-ready ensemble cell atlas of single-cell RNA and ATAC sequencing data

    Xi Xi, Yixin Chen, Xinze Wu, et al.

    Scientific Data · Dec 2025

    1
  • Weighted Diversified Sampling for Efficient Data-Driven Single-Cell Gene-Gene Interaction Discovery

    Yifan Wu, Yuntao Yang, Zirui Liu, et al.

    arXiv.org · Oct 2024

    1

Related models

Models with similar goals, methods, or subject matter.

  • CellPLM

    OmicsML

    Single-cell foundation model that treats cells as tokens and tissues as sentences, encoding cell-cell relationships across 85M parameters.

    Single-cell
  • scGPT

    Bowang Lab

    Generative pretrained transformer trained on 33 million human cells for single-cell annotation, batch correction, and perturbation prediction.

    Single-cell
  • scMOBA

    Chinese Academy of Sciences / Shanghai AI Laboratory

    Conversational single-cell and spatial multi-omics brain foundation model, with zero-shot cell annotation and disease prediction across species.

    Single-cellLanguage model
  • Cell2Sentence

    Yale University

    Framework turning single-cell expression profiles into ranked gene-name sequences, letting off-the-shelf language models generate and annotate cells.

    Single-cell
  • scFoundation

    Biomap Research

    Single-cell transcriptomics foundation model with 100 million parameters, pretrained on over 50 million human scRNA-seq profiles for cell embeddings.

    Single-cell
  • scLinguist

    Central South University

    Single-cell foundation model with a Hyena backbone that translates across omics layers, predicting protein abundance from transcriptomes zero-shot.

    Single-cell

Citations

Total Citations6
Influential0
References0

GitHub

Stars62
Forks5
Open Issues13
Contributors2
Last Push2y ago
LanguageJupyter Notebook
LicenseMIT

Fields of citing research

  • Biology100%
  • Computer Science100%
  • Medicine67%

Share of papers citing this model.

Openness

bio.rodeo opennessOpen weights · open weights, closed recipe
48Partial
Usability — can I run it?72
Reproducibility — can I retrain it?32
Model Openness Framework
Unclassified
Restrictive license on core components

Tags

foundation_modelgenerativelanguage_modelmulti_task

Resources

GitHub RepositoryResearch PaperResearch PaperDataset