Hyperbolic protein language model for alignment-free phylogenetic inference, turning ESM2-650M embeddings into distance matrices for tree placement.
Zhejiang University / Alibaba Cloud / Nanjing University of Posts and Telecommunications / Zhejiang University School of Medicine
Released March 6, 2025
Cross-domain molecular foundation model encoding small molecules, protein pockets, and their complexes in 2D and 3D on one Transformer backbone.
Alibaba Cloud / Beijing Zhongguancun Academy / Zhongguancun Institute of Artificial Intelligence / University of Science and Technology of China / Agricultural Genomics Institute at Shenzhen / Hong Kong University of Science and Technology (Guangzhou) / Hong Kong University of Science and Technology / Mila / Université de Montréal / HEC Montréal / Carnegie Mellon University
Released February 11, 2025
Long-context generative genomic foundation model with a 98k-nucleotide window, trained on 386 billion bases of eukaryotic DNA for sequence design.
Genomic foundation model for metagenomic annotation: a 500M-parameter bidirectional encoder calling coding regions at single-nucleotide resolution.
Zhejiang University / Alibaba Cloud / University of Science and Technology of China / Hong Kong University of Science and Technology (Guangzhou)
Released December 28, 2024
Protein-text foundation model aligning sequences with function descriptions through segment-wise objectives for static and dynamic functional sites.
Zhejiang University / Alibaba Cloud / University of Science and Technology of China / Liangzhu Laboratory / Second Affiliated Hospital, Zhejiang University School of Medicine
Released November 4, 2024
Protein inverse folding as a generative Markov bridge, refining a structure-derived sequence prior with a frozen protein language model.
Alibaba Cloud / Sun Yat-sen University / University of Sydney / Fudan University / Zhejiang University / Chinese Academy of Medical Sciences / Peking Union Medical College / City University of Hong Kong
Released May 10, 2024
Unified DNA, RNA, and protein foundation model with 1.8B parameters, pretrained across 169,861 species to learn the central dogma from sequence.