Chinese medical vision-language model pairing a Vision Transformer with an LLM to caption medical images and answer clinical questions in Chinese.
Fully 3D promptable segmentation foundation model for volumetric CT and MR, encoding whole volumes so anatomy can be segmented from one prompt point.
Chest X-ray vision-language model that generates free-text radiology reports, pairing a CXR-specific image encoder with a 7B LLaMA-2 language model.
Encoder-decoder framework unifying molecules, proteins, and natural language with SELFIES notation for cross-modal drug discovery tasks.
DNA language model for variant effect prediction across coding and non-coding regions, using whole-genome alignments of 100 vertebrate species.
Structure-aware protein language model pairing amino acid tokens with Foldseek 3Di structural states, outperforming ESM-2 across 10 downstream tasks.
Abdominal CT segmentation model driven by CLIP text embeddings, covering 25 organs and 6 tumor types with zero-shot extension to new categories.
Self-supervised foundation model for retinal imaging, pretrained on 1.6 million unlabelled fundus and OCT scans to detect ocular and systemic disease.
fMRI foundation model pretrained with masked autoencoding on roughly 6,700 hours of recordings for clinical prediction and network discovery.
Vision-language foundation model for pathology, fine-tuned from CLIP on 208,414 image-text pairs for zero-shot classification and image retrieval.
Radiology foundation model that reads interleaved 2D and 3D scans with text for diagnosis, visual question answering, and report generation.
Multimodal medical vision-language model for few-shot visual question answering, learning new imaging tasks from in-context examples at inference.
Multi-language transformer framework using five pre-trained language models to predict DNA methylation (6mA, 4mC, 5hmC) across species.
Genomic foundation model built on the Hyena operator, processing DNA at single-nucleotide resolution with context windows up to 1 million tokens.
Multi-species genomic foundation model swapping k-mer tokenization for byte pair encoding, matching Nucleotide Transformer with 21x fewer parameters.
Vision Transformer for cell instance segmentation and classification in H&E whole-slide images, extended by CellViT++ with foundation backbones.
Family of transformer-based DNA language models using BPE tokenization and BigBird sparse attention to reach context lengths up to 36,000 base pairs.
Biomedical vision-language assistant for question answering on radiology and pathology images, adapted from LLaVA on PubMed Central captions.
Single-cell foundation model pretrained on about 30 million human transcriptomes, using rank-value encoding for context-aware gene network inference.
Multimodal pathology assistant that answers questions about histology and cytology images, pairing the PathCLIP vision encoder with a Vicuna-13B LLM.
Generative medical visual question answering model that pairs a vision encoder with a language model, trained on the 227k-pair PMC-VQA dataset.