A global technology corporation building software, cloud, and AI, whose research reaches into biology, health, and scientific discovery.
The research division of Microsoft, spanning AI, systems, and quantum computing across global labs, with sustained work in biomedicine and health.
26 models
A Microsoft Research lab applying large-scale deep learning to scientific discovery, from molecular and drug design to new materials.
8 models
The Asia-Pacific lab of Microsoft Research, working on AI systems, large language models, and their application to health and the sciences.
9 models
Huashan Hospital, Fudan University / Fudan University / Renji Hospital, Shanghai Jiao Tong University School of Medicine / Microsoft
Released August 12, 2026
Single-cell metabolome inference from scRNA-seq, learned from spatially paired Visium and MALDI-MSI sections by multiple-instance learning.
Zhongshan Hospital, Fudan University / Fudan University / Huadong Hospital, Fudan University / Tongji University / Sir Run Run Shaw Hospital / Microsoft Research Asia / Jinling Hospital, Nanjing University Medical School / Nanjing University / The People's Hospital of Lincang
Released August 4, 2026
Renal tumor histopathology model that detects tissue regions, classifies nine subtypes, grades nuclei and scores prognosis from a single H&E slide.
Spatial proteomics prediction from routine H&E slides, generating 21-channel virtual multiplex immunofluorescence maps of the tumor microenvironment.
Distilled whole-slide pathology foundation model pairing a 22M-parameter ViT-S tile encoder with a LongNet slide encoder for cohort-scale analysis.
Generative microscopy foundation model that synthesizes in-silico fluorescence images of protein subcellular localization from amino-acid sequence.
Generative language model for phenotype-driven drug discovery, proposing small-molecule structures from up- and down-regulated gene signatures.
Microsoft Research / German Cancer Research Center (DKFZ) / University of Cambridge / Heidelberg University / Mayo Clinic
Released October 20, 2025
Vision-language encoders for chest CT that align 3D volumes with radiology reports using contrastive, report-generation, and masked-image objectives.
Protein language model conditioned on ensembles of computed conformations, giving state-aware embeddings for interaction, localization, and function.
Fudan University / Microsoft Research Asia / Shandong University / Shanghai Jiao Tong University / Zhejiang University School of Medicine / Shandong First Medical University / Linyi People's Hospital
Released August 22, 2025
Vision-language foundation model for kidney cancer CT, covering zero-shot malignancy diagnosis, report generation, and recurrence risk prediction.
China-Japan Friendship Hospital / Xidian University / Microsoft Research Asia / Peking University / Fudan University / Arizona State University
Released August 17, 2025
Dermatology foundation model pretrained on 432,776 skin images, covering malignancy classification, severity grading, and lesion segmentation.
Institute for Protein Design / University of Washington / University of Cambridge / University of Oxford / UT Southwestern Medical Center / Technical University of Denmark / Microsoft / NVIDIA / Howard Hughes Medical Institute
Released August 14, 2025
All-atom structure prediction for arbitrary biomolecular complexes of proteins, nucleic acids, and ligands, with code and weights under a BSD license.
Protein language models trained on billions of natural and synthetic sequences for de novo design and zero-shot mutation-effect prediction.
Paige AI / Microsoft Research / Memorial Sloan Kettering Cancer Center / Yale University
Released June 16, 2025
Multimodal slide-level pathology foundation model trained by clinical-dialogue supervision on 2.3M whole-slide images and 14M Q&A pairs.
Ligand-aware protein language model that cross-attends SaProt embeddings to ligand SMILES, beating SaProt across six downstream benchmarks.
Tsinghua University / Microsoft Research AI for Science / McGill University / Mila
Released March 26, 2025
Structure-based drug design model pairing an autoregressive transformer for ligand graphs with a diffusion head for 3D binding-pose coordinates.
Microsoft Research AI for Science / Peking University / Hong Kong University of Science and Technology (Guangzhou) / Beijing Institute of Mathematical Sciences and Applications / Huazhong University of Science and Technology / Tsinghua University / Microsoft Research Asia
Released March 9, 2025
Generative foundation model that co-generates sequence and 3D coordinates for proteins, small molecules, and crystals under functional objectives.
Generative design of protease substrates, producing 10-mer peptides conditioned on a target cleavage profile across 18 matrix metalloproteinases.
Chest X-ray vision-language model that drafts the findings section of a radiology report, at 7B parameters small enough to run on a single GPU.
Long-range DNA language model interleaving attention with Mamba2 state-space layers to read 131kb of sequence at single-nucleotide resolution.
Unified science foundation model treating molecules, proteins, RNA, DNA, and materials as one sequence language, in 1B, 8B, and 46.7B sizes.
Tsinghua University / Microsoft Research AI for Science / McGill University / Mila
Released February 7, 2025
Structure-based drug discovery transformer that handles protein-ligand docking and pocket-aware 3D molecule design in one pretrained model.
Microsoft Research / Imperial College London / Vector Institute / University Health Network
Released February 5, 2025
Autoregressive genomic foundation models from 20M to 1B parameters that solve ten DNA tasks at once and map sequences to text and images.
Generative model that emulates protein equilibrium ensembles, sampling cryptic pockets and unfolded states far faster than molecular dynamics.
Biomedical imaging foundation model that segments, detects, and recognizes structures across nine modalities from natural language prompts.
Protein language model that captures short- and long-range residue co-evolution through a dual pre-training objective, at 3B parameters.
Microsoft Research AI for Science / University of Washington / Tsinghua University / Beijing Normal University
Released October 27, 2024
Single-cell model that ranks the genes driving a cell state transition, using a gene graph-enhanced manifold pretrained on 20 million cells.
Multimodal protein function annotation that scores a sequence against free-text descriptions, including GO and EC labels unseen during training.
Mixed-modal DNA, RNA, and protein foundation model at 110M and 270M parameters, with in-context learning across sequence modalities.
Microsoft / Microsoft Research / University of Wisconsin-Madison / University of Washington
Released October 9, 2024
Medical imaging embedding model spanning X-ray, CT, MRI, dermoscopy, OCT, fundus, ultrasound, histopathology and mammography in one encoder.
McGill University / Shanghai Jiao Tong University / Mila / Université de Montréal / Hong Kong University of Science and Technology / Institute for Protein Design / Microsoft Research / Google DeepMind
Released October 1, 2024
Enzyme catalytic pocket design conditioned on a reaction: substrate and product in, pocket backbone, sequence, and EC class out.
Multi-task EEG foundation model that treats brain signals as a foreign language, pairing a text-aligned neural tokenizer with a GPT-2 backbone.
Text-to-text biological language model spanning molecules, proteins, and text, adding IUPAC names and multi-task instruction tuning to BioT5.
Lightweight AlphaFlow variant that fine-tunes only AlphaFold's structure module, keeping the Evoformer frozen to cut conformational sampling cost.
Microsoft Research multimodal LLM for grounded chest X-ray report generation, localizing each described finding with bounding boxes on the image.
Whole-slide histopathology foundation model pretrained on 1.3 billion image tiles from 171,189 clinical slides spanning 31 tissue types.
Deep learning framework predicting equilibrium distributions of molecular systems, enabling efficient ensemble generation and conformation sampling.
EEG foundation model pretrained with vector-quantized self-supervision, yielding interpretable discrete codes that transfer to seizure detection.
EEG pretraining framework mapping any electrode montage to a unified topology for topology-agnostic representations that transfer across datasets.
Radiology-specific multimodal LLM that generates the findings section of a chest X-ray report from a frontal image, pairing RAD-DINO with Vicuna-7B.
Microsoft Research AI for Science / MIT CSAIL / University of Oxford / University of Cambridge
Released October 8, 2023
De novo protein backbone generation by SE(3) flow matching, with motif-scaffolding built in. Samples a designable backbone in seconds on one GPU.
Microsoft Research / Microsoft Research AI for Science / University of Toronto / Stanford University
Released September 12, 2023
Discrete diffusion model for protein sequence and MSA generation, enabling controllable de novo design directly in sequence space without structure.
Antibody CDR design framework pairing a pretrained antibody language model with a hierarchical graph neural network for one-shot CDR generation.
Biomedical vision-language assistant for question answering on radiology and pathology images, adapted from LLaVA on PubMed Central captions.
Biomedical vision-language model trained contrastively on 15M PubMed Central figure-caption pairs for zero-shot classification, retrieval, and VQA.
Generative transformer pretrained on PubMed abstracts for biomedical text generation and mining, including relation extraction and question answering.
Microsoft Research Asia / Nanjing University / University of Science and Technology of China
Released September 22, 2022
Protein language model reading each residue alongside an unsupervised local-fragment token, so one encoder serves residue- and chain-level tasks.
Protein language model family built on CNNs rather than transformers, matching transformer quality while scaling linearly with sequence length.
Biomedical language model pretrained from scratch on PubMed abstracts with a WordPiece vocabulary derived from biomedical text rather than the web.