Labs & Groups (3)
Microsoft Research
The research division of Microsoft, spanning AI, systems, and quantum computing across global labs, with sustained work in biomedicine and health.
17 models
Microsoft Research AI for Science
A Microsoft Research lab applying large-scale deep learning to scientific discovery, from molecular and drug design to new materials.
1 model
Microsoft Research Asia
The Asia-Pacific lab of Microsoft Research, working on AI systems, large language models, and their application to health and the sciences.
3 models
Models (25)
Generative microscopy foundation model that synthesizes in-silico fluorescence images of protein subcellular localization from amino-acid sequence.
Generative language model for phenotype-driven drug discovery, proposing small-molecule structures from up- and down-regulated gene signatures.
CoLiPRI
Microsoft Research / German Cancer Research Center (DKFZ) / University of Cambridge / Heidelberg University / Mayo Clinic
Released October 20, 2025
Vision-language encoders for chest CT that align 3D volumes with radiology reports using contrastive, report-generation, and masked-image objectives.
Protein language model conditioned on ensembles of computed conformations, giving state-aware embeddings for interaction, localization, and function.
Protein language models trained on billions of natural and synthetic sequences for de novo design and zero-shot mutation-effect prediction.
Chest X-ray vision-language model that drafts the findings section of a radiology report, at 7B parameters small enough to run on a single GPU.
Unified science foundation model treating molecules, proteins, RNA, DNA, and materials as one sequence language, in 1B, 8B, and 46.7B sizes.
Generative model that emulates protein equilibrium ensembles, sampling cryptic pockets and unfolded states far faster than molecular dynamics.
Biomedical imaging foundation model that segments, detects, and recognizes structures across nine modalities from natural language prompts.
Protein language model that captures short- and long-range residue co-evolution through a dual pre-training objective, at 3B parameters.
Multi-task EEG foundation model that treats brain signals as a foreign language, pairing a text-aligned neural tokenizer with a GPT-2 backbone.
Text-to-text biological language model spanning molecules, proteins, and text, adding IUPAC names and multi-task instruction tuning to BioT5.
Lightweight AlphaFlow variant that fine-tunes only AlphaFold's structure module, keeping the Evoformer frozen to cut conformational sampling cost.
Microsoft Research multimodal LLM for grounded chest X-ray report generation, localizing each described finding with bounding boxes on the image.
Whole-slide histopathology foundation model pretrained on 1.3 billion image tiles from 171,189 clinical slides spanning 31 tissue types.
Deep learning framework predicting equilibrium distributions of molecular systems, enabling efficient ensemble generation and conformation sampling.
EEG foundation model pretrained with vector-quantized self-supervision, yielding interpretable discrete codes that transfer to seizure detection.
EEG pretraining framework mapping any electrode montage to a unified topology for topology-agnostic representations that transfer across datasets.
Radiology-specific multimodal LLM that generates the findings section of a chest X-ray report from a frontal image, pairing RAD-DINO with Vicuna-7B.
Discrete diffusion model for protein sequence and MSA generation, enabling controllable de novo design directly in sequence space without structure.
Antibody CDR design framework pairing a pretrained antibody language model with a hierarchical graph neural network for one-shot CDR generation.
Biomedical vision-language assistant for question answering on radiology and pathology images, adapted from LLaVA on PubMed Central captions.
Biomedical vision-language model trained contrastively on 15M PubMed Central figure-caption pairs for zero-shot classification, retrieval, and VQA.
Generative transformer pretrained on PubMed abstracts for biomedical text generation and mining, including relation extraction and question answering.
Protein language model family built on CNNs rather than transformers, matching transformer quality while scaling linearly with sequence length.