NYU Grossman School of Medicine / Leiden University Medical Center
Self-supervised colorectal histopathology model turning H&E tiles into interpretable phenotype clusters and a disease-free survival risk score.
H&E whole-slide images carry prognostic signal in tumor morphology and in the surrounding microenvironment, but that signal is hard to extract at scale and harder still to explain to the pathologists who would act on it. HPL-PanColon is a self-supervised, tile-level representation model for the colorectal adenoma–carcinoma spectrum, developed in the Division of Precision Medicine at NYU Grossman School of Medicine with Leiden University Medical Center and the UNITED collaboration, and released as a bioRxiv preprint.
It is the third model in the Histomorphological Phenotype Learning (HPL) lineage, following a lung study and a TCGA-only colon study, both published in Nature Communications. Unlike those single-cohort predecessors, HPL-PanColon is trained on a multicenter developmental cohort spanning benign adenomas through invasive colorectal cancer — the "pan-colon" in its name.
Its positioning is deliberately narrow. General-purpose pathology encoders such as UNI and TITAN are pretrained across organs and serve as default backbones for most computational-pathology pipelines. HPL-PanColon trades that breadth for organ specificity, and its central representation claim is about robustness rather than raw accuracy: its embeddings show reduced institution- and dataset-specific batch effects, the failure mode that most often undermines multicenter prognostic models.
The encoder is trained with Barlow Twins, a redundancy-reduction self-supervised objective, on roughly 245,000 H&E tiles pooled from three sources: the public TCGA-COAD cohort, the AVANT clinical-trial cohort (available under a Genentech/Roche license), and an internal NYU adenoma cohort. The training corpus therefore mixes open, restricted, and private data and cannot be reassembled from public sources alone. Tiles are embedded in 128 dimensions and clustered into histomorphological phenotype clusters covering the adenoma–carcinoma spectrum.
Evaluation uses a global survival cohort of 1,024 colorectal cancer patients from the UNITED study, an international multicenter validation of the tumor–stroma ratio in colon cancer coordinated at Leiden. With the encoder frozen, tile embeddings feed an attention-based survival model under leave-one-institution-out splits to derive CHiPS. Attention weights combined with phenotype assignments localize high risk to desmoplastic, stromal, and fibroinflammatory morphologies, and low risk to tumor-rich epithelial glandular patterns. Spatial transcriptomic analysis links the high-risk morphologies to fibroblastic, perivascular, myofibroblastic, and immune-reactive microenvironment programs, and the low-risk ones to epithelial, tumor-enriched regions.
HPL-PanColon targets prognostic stratification in colorectal cancer, where TNM staging alone leaves substantial residual risk unexplained. CHiPS provides a slide-derived score that adds information to clinicopathological variables, and because risk decomposes into named phenotype clusters, a pathologist can inspect which regions drive a prediction rather than accept an opaque number. The frozen-encoder plus refit-head recipe also gives translational groups a template for organ-specific pathology modeling on modest data volumes. All reported results are retrospective: the model is a research tool and has not been validated as a clinical diagnostic device.
The work is a counterpoint to the assumption that pan-tissue foundation models are always the right backbone. A domain-restricted encoder trained on a few hundred thousand tiles yielded representations less contaminated by site-specific signal than general-purpose alternatives, which is the property that decides whether a prognostic model survives transfer to a new institution. Coupling that encoder to interpretable phenotype clusters and then to spatial transcriptomics gives an unusually complete chain from pixels to tissue biology. Several limitations are settled: the preprint has not been peer reviewed; the released code covers the downstream survival and interpretation analyses, with the encoder built on the previously published HPL codebase and no standalone colorectal checkpoint released; and the mixed-license training corpus limits exact reproduction.
Papers that recently cited this model.
The most-cited papers that cite this model.
Providers that host HPL-PanColon for inference, fine-tuning, or weight download.
No providers recorded yet. Browse all providers
Not enough data