Breast cancer recurrence risk read from an H&E slide and six routine clinical variables, using a frozen pan-cancer pathology encoder.
No providers recorded yet. Browse all providers
Genomic assays such as Oncotype DX read prognosis out of gene expression, and the price is physical: tissue is consumed, results take ten to thirty days, and most patients who take one land in an intermediate band that does not resolve a treatment decision — 526 of the 858 assayed patients in this study's comparison cohorts. Meanwhile the H&E slide the pathologist has already read carries prognostic morphology nobody has to specify in advance: tumor architecture, stromal texture, the placement of the immune infiltrate.
Ataraxis Breast is the test built on that observation. Developed by Ataraxis AI with NYU Grossman School of Medicine and a dozen cancer centers across seven countries, it passes a digitized H&E slide from a core needle biopsy or resection through Kestrel — the company's pan-cancer pathology foundation model, a vision transformer self-supervised with DINOv2 on 400 million tissue patches — and feeds the resulting patch embeddings to supervised time-to-event models. A second score comes from six routine clinical variables, and the two are averaged into a continuous risk value between 0 and 1.
Architecturally this is a thin supervised head on a frozen general-purpose backbone, and Kestrel belongs to the same generation of DINOv2-pretrained pathology encoders as UNI and Prov-GigaPath. Where CHIEF spreads one backbone across many prediction tasks, this work aims a single head at one endpoint and spends its effort evaluating it across independent patient populations.
For the primary endpoint, disease-free interval, the pooled concordance index across the five evaluation cohorts was 0.71 [0.68–0.75], with a hazard ratio of 3.63 [3.02–4.37] per 0.2-unit increase in score. Per cohort: 0.74 in Providence (n=1,733), 0.70 in TCGA-BRCA (n=911), 0.70 in Basel (n=269), 0.67 in UChicago (n=421) and 0.62 in Karmanos (n=168). Pooled time-dependent AUC was 0.76, 0.71 and 0.73 at three, five and seven years, and the score stayed prognostic for distant recurrence-free interval (0.70) and overall survival (0.65).
Across the 858 patients also assayed with Oncotype DX, the pooled C-index was 0.67 [0.61–0.74] against 0.61 [0.49–0.73]; the journal version presents this as numerically higher discrimination, with overlapping confidence intervals. In a multivariate Cox model adjusting for Oncotype DX score, Nottingham grade, race and dataset, the AI score carried an adjusted hazard ratio of 2.95 [1.82–4.79, p < 0.001] while Oncotype DX's was not significant (1.43 [0.91–2.27, p = 0.12]). The high/low cutoff is the 80th percentile of scores in the HR+/HER2− subpopulation of the NYU training cohort, a threshold the authors call preliminary.
Because the input is a slide already cut at diagnosis, the workflow adds no wet-lab step: accessioning, scanning and inference take under an hour against ten to thirty days for a genomic assay, and the tissue block survives for later sequencing if the disease progresses. The paper cites a Medicare price of $706 for the first digital pathology procedure codes against roughly $3,800 for genomic assays. In the Oncotype-tested cohorts the model moved 666 of 858 patients into a different risk category and gave all 526 intermediate-risk patients a low or high call. Ataraxis AI sells the model commercially as Ataraxis Breast.
The work appeared as a preprint in October 2024 and in Nature Communications in 2026, and its evaluation is among the largest assembled for a breast cancer prognostic model: 8,161 women across 15 cohorts in seven countries, with more than 40% held out for external testing. The evidence is retrospective and observational, and the test is prognostic rather than predictive — it was never trained to model treatment effect, so a high score does not identify patients who will benefit from chemotherapy. The study population is women with non-metastatic invasive breast cancer, the Oncotype DX comparison is confined to HR+/HER2− patients in three of the five evaluation cohorts, and the codebase and weights are proprietary and unreleased, leaving external groups unable to compute the score independently.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.