Tsinghua University
A comprehensive research university in Beijing, pairing strength in engineering and computer science with life science and biomedical AI research.
Models (25)
Tri-modal foundation model unifying histology images, spatial transcriptomics, and language for zero-shot pathology and spatial biology reasoning.
AMix-2
Shanghai AI Laboratory / Tsinghua University / Fudan University / City University of Hong Kong / Chinese University of Hong Kong, Shenzhen
Released May 30, 2026
Protein-text foundation model placing amino acid sequences and natural language in one token space for protein understanding and de novo design.
Physics-informed generative foundation model for quantitative diffusion MRI that maps brain microstructure and adapts zero-shot to each participant.
Protein structure prediction and binder design in a single generative step, replacing AlphaFold3's iterative diffusion sampling with one forward pass.
Co-generative protein language model decoding sequence and structure tokens together from GO functional annotations for de novo protein design.
All-atom generative foundation model that designs small molecules, peptides, and nanobodies against a target binding site from a single checkpoint.
Mol-Reasoning
Pengcheng Laboratory / Sun Yat-sen University / Tsinghua University
Released March 13, 2026
Molecular reasoning model built on DeepSeek-7B, using chain-of-thought and reinforcement learning for property prediction, generation, and reactions.
Hierarchical language model for atlas-level cell-type annotation of scATAC-seq data that annotates new query datasets without retraining.
Protein language model that encodes sequences as discrete words from a learned vocabulary for zero-shot function inference and protein design.
Latent diffusion model that designs D-peptide binders against native L-protein targets, generalizing across chirality via axial vector features.
OmniNovo
Fudan University / Shanghai AI Laboratory / Tsinghua University / Westlake University / Tongji University / Shanghai Innovation Institute / University of British Columbia / Zhejiang University / Stony Brook University
Released December 13, 2025
De novo peptide sequencing transformer that reads modified and unmodified peptides directly from tandem mass spectra without a reference database.
Equivariant diffusion model that converts peptide binders into drug-like small molecules, generating peptidomimetics inside the target protein pocket.
Multimodal LLM that tokenizes single cells into discrete VQ-VAE codebook tokens, letting one model reason jointly over transcriptomes and text.
Retrieval-augmented latent diffusion model for protein binder design, retrieving interfaces in a shared latent space across peptides and antibodies.
BrainOmni
Tsinghua University / Shanghai AI Laboratory / University of Cambridge / University College London
Released May 18, 2025
Brain foundation model unifying EEG and MEG in a single encoder via a shared discrete tokenizer that transfers across sensor layouts and montages.
Hypergraph foundation model for brain disease diagnosis from resting-state fMRI, self-supervised on high-order connectivity among brain regions.
Latent diffusion model for single-cell multi-omics generation and modality translation, with gradient-based inference of gene regulatory networks.
Multimodal ECG language model pairing a specialized signal encoder with a biomedical LLM for cardiovascular disease detection and question answering.
Brainfound
Tsinghua University / Chinese PLA General Hospital / Beijing Tiantan Hospital
Released January 10, 2025
Multimodal vision-text foundation model for brain CT and MRI, pretrained on roughly 10 million image-report pairs to act as a clinical copilot.
End-to-end framework predicting protein structure and mutational fitness from a single sequence, with five-fold faster inference than ESMFold.
RNA language model that builds base-pairing constraints into self-attention, pretrained on 20.4 million sequences for structure and function tasks.
Generative language model for single-cell transcriptomics with 368M parameters, unifying cell type annotation, batch integration, and cell generation.
Unified 100-billion-parameter protein language model combining autoencoding and autoregressive objectives for protein understanding and generation.
Diffusion model for synthesizing single-cell RNA-seq data, with guided generation of specific cell types, rare cells, and developmental trajectories.
Transformer model predicting context-specific epigenomic signals across cell types using DNA sequence and transcription factor activity profiles.