Models (4)
HuatuoGPT-Vision
Shenzhen Research Institute of Big Data / Chinese University of Hong Kong, Shenzhen
Released June 27, 2024
Open medical multimodal LLMs (7B and 34B) for visual question answering over radiology, pathology, and endoscopy images, trained on PubMedVision.
PTUnifier
Chinese University of Hong Kong, Shenzhen / Sun Yat-sen University / Shenzhen Research Institute of Big Data
Released February 17, 2023
Medical vision-language pretraining unifying fusion-encoder and dual-encoder designs, handling image-only, text-only, and paired inputs in one model.
ARL (Align, Reason and Learn)
Shenzhen Research Institute of Big Data / Chinese University of Hong Kong, Shenzhen / Sun Yat-sen University
Released September 15, 2022
Medical vision-language pretraining framework that injects structured medical knowledge into radiology image-text learning for VQA and retrieval.
M3AE
Shenzhen Research Institute of Big Data / Chinese University of Hong Kong, Shenzhen / Sun Yat-sen University
Released September 15, 2022
Self-supervised medical vision-and-language pretraining via multi-modal masked autoencoders that reconstruct masked image patches and text tokens.