bio.rodeo
ModelsOrganizationsLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Imaging foundation models
Imaging

BiomedParse

Microsoft Research

Biomedical imaging foundation model that segments, detects, and recognizes structures across nine modalities from natural language prompts.

Released: November 2024

BiomedParse is a biomedical foundation model from Microsoft Research that unifies segmentation, detection, and recognition of anatomical structures and abnormalities across nine imaging modalities within a single framework. Rather than requiring manual bounding boxes or point clicks, BiomedParse accepts a natural language description of a target structure and returns a pixel-level segmentation mask — making it operable by any researcher who can describe what they are looking for. The model was published in Nature Methods in November 2024.

Most biomedical image analysis systems are modality-specific, task-specific, or require expert interaction through point and bounding-box prompts. BiomedParse addresses all three constraints simultaneously. Given a text prompt such as "liver" or "polyp", the model produces segmentation masks across CT, MRI, X-ray, pathology slides, endoscopy, ultrasound, fundus photography, dermoscopy, and optical coherence tomography (OCT) images — without any modality-specific fine-tuning.

A key insight behind the model is that training segmentation, detection, and recognition jointly produces mutual regularization: each task improves the others through shared representations. This joint learning formulation allows BiomedParse to learn more generalizable visual features than single-task baselines trained on the same data.

#Key Features

  • Nine imaging modalities in one model: CT, MRI, X-ray, pathology, endoscopy, ultrasound, fundus, dermoscopy, and OCT are all supported without modality-specific fine-tuning or separate model instances.
  • Text-only prompting: Natural language descriptions replace bounding boxes and point prompts, reducing required user operations from hundreds to one in typical cell segmentation workflows.
  • Joint task learning: Segmentation, detection, and recognition are trained simultaneously, with performance on each task improving the others through shared visual representations.
  • Invalid input detection: A Kolmogorov-Smirnov test on predicted pixel probability distributions flags queries describing objects absent from the image, achieving 0.93 precision and 1.00 recall (AUROC 0.99).
  • 82 object types across 25 anatomical sites: Covers organs, abnormalities, and histological structures, organized around a standardized biomedical taxonomy derived from 45 public datasets.
  • GPT-4 data harmonization: Heterogeneous natural-language labels from source datasets are standardized into a consistent ontology using GPT-4, enabling unified training across disparate annotation conventions.

#Technical Details

BiomedParse is built on the SEEM (Segment Everything Everywhere all at once with Multi-granularity) framework, extended with biomedical domain adaptations. The image encoder is a Focal Vision Transformer initialized from a pretrained checkpoint and fine-tuned on biomedical images. The text encoder is PubMedBERT, which provides biomedical domain-specific language representations for interpreting clinical terminology in prompts. A transformer-based mask decoder cross-attends image and text features to produce pixel-wise segmentation probability maps at the input image resolution. An auxiliary meta-object classifier is trained on 15 intermediate semantic categories (organ, abnormality, histology) to jointly train the image encoder with coarse object semantics alongside the fine-grained mask decoder.

Training data was constructed from 45 publicly available biomedical segmentation datasets and comprises 1.1 million images, 3.4 million image-mask-label triples, and 6.8 million image-mask-description triples after GPT-4 synthesis of synonymous descriptions for each label. On a held-out test set of 102,855 instances spanning all nine modalities and 64 major object types, BiomedParse achieved the highest Dice scores across all nine modalities compared to competing methods, with statistically significant improvement over MedSAM using oracle bounding boxes (p < 10^-4). On detection of irregular-shaped objects, it improved Dice by 39.6% over the best competing method. On recognition, it improved F1 by 74.5% relative to Grounding DINO.

#Applications

BiomedParse is designed for researchers and clinical informaticists who need rapid, text-driven quantification of anatomical objects across large image cohorts without modality-specific tooling. In radiology, it enables automated organ and lesion segmentation in CT and MRI without drawing bounding boxes. In pathology, it segments cells, glands, and tissue structures from whole-slide image patches using plain-language descriptions. In ophthalmology, it handles retinal structure and lesion segmentation in fundus photographs and OCT volumes. It is equally applicable to dermatology (skin lesion delineation in dermoscopy) and endoscopy (polyp and mucosal abnormality segmentation). Beyond inference, BiomedParse can accelerate expert annotation workflows by generating initial segmentation masks from text descriptions that human annotators then refine.

#Impact

BiomedParse represents a meaningful step toward general-purpose biomedical image analysis, demonstrating that a single model can match or exceed task-specific and modality-specific baselines when trained with a sufficiently diverse and well-harmonized dataset. Its release on HuggingFace under the Apache-2.0 license, alongside the full BiomedParseData training corpus, lowers the barrier for downstream fine-tuning and benchmarking. Notable limitations include 2D-only processing in the v1 model (3D volumetric inference requires slice-by-slice application), a closed object vocabulary of 82 trained types that may not generalize to novel structures, and sensitivity to the alignment between user prompts and the training ontology. The model has not undergone regulatory review and should not be used for clinical decision-making without further validation.

Citation

A foundation model for joint segmentation, detection and recognition of biomedical objects across nine modalities

Zhao, T., et al. (2024) A foundation model for joint segmentation, detection and recognition of biomedical objects across nine modalities. Nature Methods.

DOI: 10.1038/s41592-024-02499-w

Recent citations

Papers that recently cited this model.

  • Prototype-based AI triage for 3D pathology

    Renao Yan, Gan Gao, Andrew H. Song, et al.

    bioRxiv · Jul 2026

    0Influential
  • Language-Guided Segmentation of Medical Images: A Review of Foundation Models

    Saqib Qamar

    Bioengineering · Jul 2026

    0
  • Learning To Focus: Anatomy-Guided Attention Regularization for Medical Image Classification

    Tonmoy Hossain, Atiqur Rahman, Farhana Hossain Swarnali, et al.

    Jul 2026

    0Influential

Top citations

The most-cited papers that cite this model.

  • MMedAgent: Learning to Use Medical Tools with Multi-modal Agent

    Binxu Li, Tian Yan, Yuanting Pan, et al.

    Conference on Empirical Methods in Natural Language Processing · Jul 2024

    120
  • MedSAM2: Segment Anything in 3D Medical Images and Videos

    Jun Ma, Zongxin Yang, Sumin Kim, et al.

    arXiv.org · Apr 2025

    96
  • Large-vocabulary segmentation for medical images with text prompts

    Ziheng Zhao, Yao Zhang, Chaoyi Wu, et al.

    npj Digital Medicine · Dec 2023

    77Influential
  • Segment Anything in Medical Images and Videos: Benchmark and Deployment

    Jun Ma, Sumin Kim, Feifei Li, et al.

    arXiv.org · Aug 2024

    72
  • VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge

    Vishwesh Nath, Wenqi Li, Dong Yang, et al.

    Computer Vision and Pattern Recognition · Nov 2024

    61

Related models

Models with similar goals, methods, or subject matter.

  • UniBiomed

    Hong Kong University of Science and Technology / Weill Cornell Medicine / Harvard University

    Universal foundation model that jointly generates diagnostic text and segments the corresponding targets across ten biomedical imaging modalities.

    ImagingLanguage model
  • MedSAM

    Bowang Lab / University Health Network / University of Toronto / Vector Institute / Western University / New York University / Yale University

    Promptable foundation model for universal medical image segmentation, fine-tuned from SAM on 1.57M image-mask pairs across 10 imaging modalities.

    Imaging
  • BiomedGPT

    Lehigh University / University of Georgia / Stanford University / Massachusetts General Hospital / University of Pennsylvania / University of Central Florida / UC Santa Cruz / UTHealth Houston / Mayo Clinic / Samsung Research America

    Open-source, lightweight generalist vision-language foundation model for diverse biomedical imaging and text tasks.

    Language modelImagingPathology
  • SegVol

    Beijing Academy of Artificial Intelligence

    Promptable 3D foundation model for volumetric CT segmentation, covering over 200 anatomical categories through point, box, and free-text prompts.

    Imaging
  • SAM-Med2D

    Shanghai AI Laboratory

    Medical imaging adaptation of the Segment Anything Model, fine-tuned on 4.6M images and 19.7M masks for promptable segmentation across 10 modalities.

    Imaging

Citations

Total Citations170
Influential17
References79

GitHub

Stars686
Forks104
Open Issues42
Contributors5
Last Push6mo ago
LanguagePython
LicenseApache-2.0

HuggingFace

Downloads548
Likes110
Last Modified9mo ago

Fields of citing research

  • Computer Science96%
  • Medicine87%
  • Engineering31%
  • Biology11%
  • Environmental Science2%
  • Physics1%
  • Mathematics1%
  • Chemistry1%

Share of papers citing this model.

Openness

bio.rodeo opennessFully open · usable and reproducible
60Partial
Usability — can I run it?69
Reproducibility — can I retrain it?57
Model Openness Framework
Unclassified
Restrictive license on core components

Tags

foundation_modelmedical_imagingmultimodalradiologysegmentationtext_guidedvision_model

Resources

GitHub RepositoryResearch PaperResearch PaperOfficial WebsiteHuggingFace ModelDataset