Spatial virtual-cell foundation model that simulates single-cell gene expression at chosen positions in real human tumor tissue.
No providers recorded yet. Browse all providers
PD1 is the receptor the most widely used cancer immunotherapies block, and yet in a 1000-plex spatial transcriptomics run it appears as a scatter of individual transcripts — real signal, far too sparse to say where in a section a T cell would be shut down. Worse, the cell type you want to interrogate is often absent from the position you want to interrogate. OCTO-VC inverts the question: rather than read what a tissue contains, it predicts what a cell would express if placed at a given spot, conditioned on the neighbors that are really there.
OCTO-VC — also written OCTO-vc, OCTO-VirtualCell, and "OCTO-virtual cell" — is Noetik's spatial virtual-cell foundation model, introduced in a company technical report in December 2024. It is a transformer trained by masked token modeling on 1000-plex spatial transcriptomes from human tumor resections; the pretraining task is deliberately plain, reconstructing a center cell's transcriptome from those of the cells around it. Because the objective is generative and the model accepts a prompt — a short set of genes switched on or off to define a cell type — the checkpoint can be run in a mode the assay cannot support: drop a "virtual" CD8 T cell at thousands of coordinates across a real patient section and read back predicted expression for the rest of the panel.
OCTO-VC is a sibling within Noetik's OCTO family, not a version of any member of it: the parent model OCTO (Oncology Counterfactual Therapeutics Oracle) is multimodal — multiplex protein, spatial expression, exome and H&E — and emits a 16-channel protein image, while TARIO-2 runs the opposite direction, inferring transcriptome from H&E. The model is proprietary — no weights, code, or training data have been released — and the evidence base is Noetik's own technical reports, conference abstracts, and partner announcements rather than peer-reviewed publication.
Noetik describes OCTO-VC as a multi-scale, multimodal transformer trained via masked token modeling, with tokens encoding gene expression and user prompts specifying genes expressed in the virtual cell. Corpus size is reported at two dates: nearly 40 million cells from over 1,000 tumor samples in the December 2024 technical report, and 77 million cells across roughly 2,500 patients and more than a dozen cancer types in a September 2025 company post. Parameter count and context length are not disclosed. Validation is reported as qualitative agreement with held-out data and known biology: virtual B cells express the immature marker TCL1A inside tertiary lymphoid structures and the mature marker IGHA1 outside, and counterfactually raising tumor MHC I increases GZMB in nearby virtual CD8 T cells for most, but notably not all, patients.
Noetik's stated initial focus is cancer immunotherapy, where much patient-to-patient variation is spatial. Reported uses span clinical-stage stratification and preclinical target discovery: separating anti-PD-1 responders in a 39-patient PD-L1-positive cohort from unsupervised tumor embeddings; a virtual screen across KRAS/STK11-mutant lung tumors that traced depressed granzyme GZMA and GZMK levels in virtual CD8 T cells to a candidate adhesion-protein target; and, at SITC 2025, virtual "clonal neighborhood" simulations identifying fibroblast- and tumor-driven macrophage programs in NSCLC. Under a five-year agreement announced in January 2026, GSK holds a non-exclusive license to OCTO-VC in NSCLC and colorectal cancer.
OCTO-VC's clearest significance is what it trains on. Most virtual-cell efforts learn from cancer cell lines; OCTO-VC learns entirely from human tumor resections, shortening the path from an in-silico result to a clinical hypothesis. It is also a test of model licensing as a business structure: Noetik reported in August 2026 that delivering OCTO-VC inference and fine-tuning capability to GSK triggered the collaboration's first milestone under a subscription-style arrangement backed by $50 million in upfront capital and near-term milestones. The limitations are equally clear — every result is developer-reported, the validation shown is qualitative rather than benchmarked against competing methods, counterfactual knockout effects are correlational rather than causal by the authors' own framing, and weights, code, and corpus are all closed.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.