bio.rodeo
ModelsOrganizationsLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Pathology foundation models
Pathology

Hibou

HistAI

Histopathology foundation models pretrained with DINOv2 on over 1 million whole-slide images, released as Hibou-B and Hibou-L under Apache 2.0.

Released: June 2024
Parameters: 307 Million

Hibou is a family of Vision Transformer foundation models for digital pathology, developed by HistAI and released in June 2024. The family comprises two variants — Hibou-B and Hibou-L — pretrained on a curated dataset of over 1 million whole-slide images (WSIs) using the DINOv2 self-supervised learning framework with additional register tokens for improved feature quality.

What distinguishes Hibou from competing pathology foundation models is its combination of training scale, stain diversity, and permissive licensing. The pretraining corpus spans both H&E-stained slides (936,441 WSIs) and non-H&E modalities (202,464 slides including immunohistochemistry, special stains, and cytology), exposing the model to the full breadth of tissue preparation techniques encountered in real clinical and research settings. Both variants are released under the Apache 2.0 license, enabling unrestricted commercial and research use without the restrictive gating common among competing pathology foundation models such as Prov-GigaPath and Virchow.

At time of publication, Hibou-L established state-of-the-art average accuracy across six standard patch classification datasets and outperformed Prov-GigaPath on all three slide-level WSI classification benchmarks evaluated. Hibou-B, despite having roughly 13 times fewer parameters than GigaPath, matched or exceeded it on two of three slide-level tasks, demonstrating strong parameter efficiency from the DINOv2 training strategy.

#Key Features

  • Two model sizes: Hibou-B (ViT-B/14, ~86M parameters) and Hibou-L (ViT-L/14, ~307M parameters) accommodate different compute budgets without sacrificing representational quality.
  • DINOv2 with register tokens: Self-supervised pretraining is extended with learnable register tokens that improve attention map quality and reduce artifacts, yielding cleaner patch-level features than standard DINOv2.
  • Multi-stain training corpus: Coverage of H&E, immunohistochemistry, and special stains ensures features generalize across staining protocols, unlike models trained exclusively on H&E slides.
  • Apache 2.0 license: Fully permissive for commercial and research use, with no gating or institutional registration requirements beyond a HuggingFace account.
  • HuggingFace integration: Both variants load directly via the transformers library with a single AutoModel.from_pretrained call, simplifying integration into existing PyTorch pipelines.
  • State-of-the-art slide-level benchmarks: Hibou-L outperforms Prov-GigaPath on TCGA-BRCA, TCGA-NSCLC, and TCGA-RCC WSI classification tasks using attention-based multiple instance learning pooling.

#Technical Details

Both Hibou variants are built on the DINOv2 Vision Transformer architecture with a modification to incorporate register tokens — additional learnable tokens appended to the patch sequence that allow the model to offload global information processing away from local patch tokens, improving spatial feature quality. Hibou-B uses a ViT-B/14 backbone (85.7M parameters, 14-pixel patch size) and Hibou-L uses a ViT-L/14 backbone (~307M parameters, 14-pixel patch size). The choice of 14-pixel rather than the more common 16-pixel patch size yields finer spatial resolution per token, which is advantageous for pathology images where cellular-level features at high magnification are diagnostically relevant.

The pretraining corpus totaled over 1.1 million WSIs: 936,441 H&E slides, 202,464 non-H&E slides, and 2,676 cytology slides, sourced from public and proprietary collections covering multiple human organ systems. Hibou-L trained on approximately 1.2 billion clean patches over 1.175 million iterations on 32 NVIDIA A100-40G GPUs; Hibou-B trained on 512 million patches over 500,000 iterations on 8 A100-80G GPUs. Standard DINOv2 solarization augmentation was deliberately excluded, as it degrades performance on stained tissue images; instead, RandStainNA stain normalization and color jittering were applied. On patch classification benchmarks using linear probing, Hibou-L achieved an average accuracy of 0.890 across six datasets (CRC-100K, PCAM, MHIST, MSI-CRC, MSI-STAD, TIL-DET), surpassing contemporaneous models including Phikon, Kaiko-B8, Virchow, RudolfV, Prov-GigaPath, and H-optimus-0.

#Applications

Hibou functions as a general-purpose feature extractor for digital pathology workflows. Downstream tasks include cancer subtyping from WSI patches (e.g., distinguishing IDC from ILC in breast cancer, or LUAD from LUSC in lung cancer), molecular biomarker prediction from H&E slides (microsatellite instability, mutation status), and survival analysis using slide-level aggregated embeddings. The companion CellViT-Hibou-L model — combining Hibou-L features with the CellViT segmentation framework — enables panoptic nuclei segmentation on the PanNuke benchmark, with improved performance over CellViT-SAM-H baselines for epithelial and dead cell categories. Because Hibou was pretrained on non-H&E stains, its representations transfer more reliably to IHC panels and special stain workflows than models trained exclusively on H&E, broadening applicability across clinical laboratory settings.

#Impact

Hibou addresses a recognized gap in the pathology foundation model landscape: the combination of open licensing, multi-stain pretraining, and competitive benchmark performance has made it one of the more practically accessible models in the field. Its Apache 2.0 release stands in contrast to the non-commercial or gated licensing of several higher-profile competitors, lowering barriers for both academic research and clinical product development. A notable limitation is that the pretraining WSI dataset is not publicly released, limiting reproducibility of the pretraining procedure. Hibou-L was also trained on approximately one-sixth of HistAI's full proprietary dataset at time of publication, suggesting meaningful headroom for further performance improvement. As with all pathology foundation models, downstream applications require independent clinical validation before deployment in regulated healthcare settings.

Citation

Hibou: A Family of Foundational Vision Transformers for Pathology

Preprint

Nechaev, D., et al. (2024) Hibou: A Family of Foundational Vision Transformers for Pathology. arXiv.org.

DOI: 10.48550/arXiv.2406.05074

Recent citations

Papers that recently cited this model.

  • The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models

    D. Yajnik, Amina Asif, F. Minhas

    Jul 2026

    0
  • CellPrior-Net: Prior-Guided Nuclei Detection and Classification for H&E Whole-Slide Images

    Falah Jabar, Pasquale Lombardi, Aria Torkpour, et al.

    Jul 2026

    0
  • Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning

    Gbègninougbo Aurel Davy Tchokponhoue, Sevda Ougut, Ali Idri, et al.

    Jun 2026

    0

Top citations

The most-cited papers that cite this model.

  • Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology

    Eric Zimmermann, E. Vorontsov, Julian Viret, et al.

    arXiv.org · Aug 2024

    202
  • Artificial intelligence in digital pathology — time for a reality check

    Arpit Aggarwal, Satvika Bharadwaj, Germán Corredor, et al.

    Nature Reviews Clinical Oncology · Feb 2025

    46
  • A comprehensive evaluation of histopathology foundation models for ovarian cancer subtype classification

    Jack Breen, Katie Allen, K. Zucker, et al.

    npj Precision Oncology · May 2024

    41Influential
  • Distilling foundation models for robust and efficient models in digital pathology

    Alexandre Filiot, N. Dop, Oussama Tchita, et al.

    International Conference on Medical Image Computing and Computer-Assisted Intervention · Jan 2025

    28
  • PathBench: A comprehensive comparison benchmark for pathology foundation models towards precision oncology

    Jiabo Ma, Yingxue Xu, Fengtao Zhou, et al.

    arXiv.org · May 2025

    26

Related models

Models with similar goals, methods, or subject matter.

  • H-optimus-0

    Bioptimus

    Histopathology vision transformer with 1.1B parameters, pretrained on patches from 500,000 H&E whole-slide images across 4,000 clinical practices.

    Pathology
  • Virchow

    Paige AI

    Histopathology foundation models: self-supervised vision transformers pretrained on millions of whole-slide images for tile-level feature extraction.

    Pathology
  • Prov-GigaPath

    Microsoft Research

    Whole-slide histopathology foundation model pretrained on 1.3 billion image tiles from 171,189 clinical slides spanning 31 tissue types.

    Pathology
  • GenBio-PathFM

    genbio.ai

    Histopathology foundation model with 1.1B parameters, trained entirely on public data using JEDI, a dual-stage strategy combining JEPA and DINO.

    Pathology
  • UNI

    Mahmood Lab

    Computational pathology foundation model (ViT-L/16, DINOv2) pretrained on over 100 million H&E tiles from more than 100,000 whole-slide images.

    Pathology
  • Phikon

    Owkin

    Self-supervised histopathology foundation models for H&E whole-slide images. Phikon-v2 is a ViT-L/16 DINOv2 encoder used for biomarker prediction.

    Pathology

Citations

Total Citations91
Influential8
References16

GitHub

Stars79
Forks7
Open Issues1
Contributors1
Last Push1y ago
LanguagePython
LicenseApache-2.0

HuggingFace

Downloads34.2K
Likes21
Last Modified1y ago
Pipelineimage-feature-extraction

Fields of citing research

  • Medicine98%
  • Computer Science97%
  • Engineering17%
  • Biology13%

Share of papers citing this model.

Openness

bio.rodeo opennessOpen weights · open weights, closed recipe
46Partial
Usability — can I run it?87
Reproducibility — can I retrain it?4
open weights, closed recipe
Model Openness Framework
Unclassified
Restrictive license on core components

Tags

foundation_modelhistologyself_supervisedvision_model

Resources

GitHub RepositoryResearch PaperOfficial WebsiteHuggingFace ModelHuggingFace Model