Every biological foundation model, evaluated and ranked by the bio.rodeo team
Showing 193–216 of 221 filtered models
EHR foundation model that writes patient histories as token sequences carrying explicit visit and day-interval tokens, invertible back to OMOP tables.
Instruction-tuned vision-language foundation model for chest X-ray interpretation, with 8 billion parameters spanning eight clinical task types.
Vision-language pretraining for 3D CT volumes, aligning scans with their radiology reports for zero-shot classification, retrieval, and segmentation.
Radiology-specific multimodal LLM that generates the findings section of a chest X-ray report from a frontal image, pairing RAD-DINO with Vicuna-7B.
Chinese medical vision-language model pairing a Vision Transformer with an LLM to caption medical images and answer clinical questions in Chinese.
Chest X-ray vision-language model that generates free-text radiology reports, pairing a CXR-specific image encoder with a 7B LLaMA-2 language model.
Encoder-decoder framework unifying molecules, proteins, and natural language with SELFIES notation for cross-modal drug discovery tasks.
Radiology foundation model that reads interleaved 2D and 3D scans with text for diagnosis, visual question answering, and report generation.
Open large language models for natural science, fine-tuned on physics, chemistry, and materials science literature with automated instruction tuning.
Multimodal medical vision-language model for few-shot visual question answering, learning new imaging tasks from in-context examples at inference.
Google's generalist multimodal biomedical AI that encodes clinical text, medical images, and genomics with a single set of weights across 14 tasks.
Biomedical vision-language assistant for question answering on radiology and pathology images, adapted from LLaVA on PubMed Central captions.
Multimodal pathology assistant that answers questions about histology and cytology images, pairing the PathCLIP vision encoder with a Vicuna-13B LLM.
Multi-modal LLM answering free-form questions about a compound's indications, pharmacodynamics and mechanism of action from its SMILES string.
Vision-language framework for 3D medical image diagnosis and visual question answering, bridging frozen image encoders and LLMs, shown on brain MRI.
Generative medical visual question answering model that pairs a vision encoder with a language model, trained on the 227k-pair PMC-VQA dataset.
Drug pair synergy prediction for rare cancer tissues, read from a language model's representation of a screening row written out as a sentence.
Medical vision-language pretraining unifying fusion-encoder and dual-encoder designs, handling image-only, text-only, and paired inputs in one model.
Text-conditioned latent diffusion model that generates synthetic chest X-rays from free-form radiology prompts by adapting Stable Diffusion.
Generative transformer pretrained on PubMed abstracts for biomedical text generation and mining, including relation extraction and question answering.
Medical vision-language pretraining framework that injects structured medical knowledge into radiology image-text learning for VQA and retrieval.
Self-supervised medical vision-and-language pretraining via multi-modal masked autoencoders that reconstruct masked image patches and text tokens.
Medical-domain CLIP fine-tuned on radiology image-caption pairs from ROCO, serving as a drop-in visual encoder for medical visual question answering.