Lingang Laboratory / Shanghai Academy of Artificial Intelligence for Science / Shanghai Institute of Biochemistry and Cell Biology / Shanghai Institute of Nutrition and Health, Chinese Academy of Sciences / Institute of Neuroscience, Chinese Academy of Sciences / Shanghai Jiao Tong University / University of Chinese Academy of Sciences / University of Tokyo
Cross-species brain spatial transcriptomics foundation model pretrained on 133M cells from human, macaque, marmoset, and mouse whole brains.
Single-cell spatial transcriptomics can map gene expression across an entire brain while preserving where each cell sits, but the resulting atlases are fragmented. Different species, anatomical coverage, and assay chemistries produce datasets that cannot be compared directly, so most analyses stop at the boundary of a single study. BrainBeacon is a foundation model built to dissolve those boundaries: it is pretrained on BrainST-133M, a corpus of 133 million spatially resolved cells covering 210,194 mm² of tissue from human, macaque, marmoset, and mouse whole brains, profiled across five platforms — MERFISH, STARmap, Xenium, Slide-seqV2, and Stereo-seq.
The model was developed by Lingang Laboratory with collaborators at several Chinese Academy of Sciences institutes in Shanghai, Shanghai Jiao Tong University, and the University of Tokyo, and posted as a bioRxiv preprint in July 2025. Where earlier single-cell foundation models such as scGPT and CellPLM learn primarily from dissociated profiles, and Nicheformer transfers spatial context between dissociated and spatial assays, BrainBeacon is trained end to end on whole-brain spatial atlases and treats evolutionary correspondence as part of the representation itself. Genes are mapped through one-to-one ortholog tables so that a cell from any of the four species is tokenized into a single shared vocabulary.
The outcome is one fixed pretrained checkpoint whose embeddings support spatial clustering with no task-specific training, and which can be fine-tuned for cell-type annotation, cross-species label transfer, and in-silico perturbation of both genes and cellular niches. The work remains a preprint awaiting peer review.
BrainBeacon's encoder is an OmicsFormer-style masked-modeling transformer with a variational latent bottleneck, adapted from the CellPLM implementation and extended for spatial and cross-species inputs. Pretraining proceeds in two stages over BrainST-133M, first over gene-expression rankings within individual cells and then over the spatial arrangement of cells relative to their neighbors, with evolutionary gene relationships supplied as a fixed prior throughout. The distributed pretraining checkpoint is a 0.33-billion-parameter model, released alongside a fine-tuned cellformer checkpoint. Reference resources bundled with the code include the ESM-2 gene embedding matrix, Ensembl-derived ortholog tables for all four species, and gene-wise expression statistics per platform.
Downstream evaluations released with the model cover spatial clustering and zero-shot cell-type annotation on a human brain MERFISH atlas, label transfer from macaque single-nucleus RNA-seq to macaque spatial data and from macaque to human spatial data, and both perturbation families. Code is MIT-licensed; the pretrained weights are distributed through a cloud storage folder rather than a model hub, and the repository ships tutorial notebooks plus hosted API documentation.
The model is aimed at neuroscientists and computational biologists working with whole-brain spatial atlases. Practical uses include annotating cell types in a new spatial dataset without a matched reference, transferring hard-won human or macaque annotations onto a newly profiled species or platform, harmonizing atlases built with incompatible gene panels, and constructing virtual brain atlases that span species. The perturbation pipelines let groups screen candidate genes and niche compositions computationally — asking how knocking out a gene in one cell type reshapes the transcriptional state of its neighbors — before committing to animal experiments. The paper applies this to spatial regulatory mechanisms of aging.
BrainBeacon extends the foundation-model approach from dissociated single cells into whole-brain spatial data and makes cross-species comparison part of pretraining rather than a post-hoc alignment step. That framing matters for a field that relies on model organisms to reason about human brains: a shared embedding space gives a quantitative handle on which cellular architectures are conserved between mouse, marmoset, macaque, and human and which are not. Adoption is early — the repository has attracted little community activity so far, and reported results have not been independently reproduced. Practical friction remains as well: inference requires downloading several large prior-knowledge files and checkpoints from separate hosts, and no model card or dataset card accompanies the release.
Papers that recently cited this model.
The most-cited papers that cite this model.
Providers that host BrainBeacon for inference, fine-tuning, or weight download.
No providers recorded yet. Browse all providers
Not enough data