bio.rodeo
ModelsOrganizationsLeaderboardAbout
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
DNA & Gene foundation models
DNA & GeneLanguage model

VirtualCRISPR

Chan Zuckerberg Biohub Chicago / Northwestern University / University of Chicago

Large language model trained on functional genomics data to prioritize novel therapeutic targets from genome-wide CRISPR knockout screens.

Released: February 2026

VirtualCRISPR is a large language model trained on functional genomics data and used to prioritize therapeutically tractable targets emerging from genome-wide CRISPR knockout screens. Rather than editing DNA or designing guide RNAs, it operates one step downstream: given the thousands of gene-level hits a pooled screen produces, it scores which of those hits represent genuinely novel biology worth pursuing as drug targets. It was developed by researchers at Chan Zuckerberg Biohub Chicago and Northwestern University and introduced in a bioRxiv preprint in February 2026.

CRISPR knockout screens can implicate hundreds to thousands of genes in a phenotype, but the bottleneck in translating a screen into a therapeutic program is triage: distinguishing hits that merely re-confirm well-studied biology from those that point to unexplored, druggable mechanisms. VirtualCRISPR addresses this triage problem. Applied as a fixed pretrained model to a completely new screen, it computes a per-gene assessment that separates established gene-phenotype relationships from underexplored ones, acting as a novelty filter that surfaces high-confidence hits with minimal prior association to the disease under study.

This places VirtualCRISPR in a different niche from other CRISPR-focused models in the catalog. Where models such as OpenCRISPR-1 generate Cas nucleases and crisprSFM predicts guide-target specificity, VirtualCRISPR reasons over the biological output of a screen to guide target selection, positioning the language model as a target-discovery engine layered on top of functional genomics.

#Key Features

  • Zero-shot novelty filtering: The pretrained model is applied without retraining to a new, unseen genome-wide screen, scoring each gene by how established versus novel its link to the measured phenotype is.
  • Trained on functional genomics data: VirtualCRISPR learns from functional genomics data, giving it a prior over gene-phenotype relationships that it uses to flag underexplored hits.
  • Screen-scale prioritization: It ranks candidates across the full set of screened genes rather than a hand-picked shortlist, compressing thousands of hits into a small number of high-confidence, tractable targets.
  • End-to-end validation: In its introducing study, the model's top novel picks were carried through molecular, cellular, and in vivo validation, demonstrating that its prioritizations translate into experimentally confirmed biology.

#Technical Details

VirtualCRISPR is a large language model trained on functional genomics data and applied as a fixed pretrained scorer to a novel dataset. That dataset was the first genome-wide CRISPR knockout screen in primary human adult epidermal keratinocytes: a library of roughly 77,000 guide RNAs was used to knock out approximately 19,000 genes in cells from two donors, with the readout being each gene's effect on IL-17 receptor A (IL17RA), a central node in psoriatic inflammation. From those 19,000+ screened genes, VirtualCRISPR prioritized arachidonate 5-lipoxygenase (ALOX5) and oxytocin receptor (OXTR) as high-confidence novel hits carrying minimal prior association with psoriasis. The preprint does not disclose the model's parameter count or detailed architecture, and no public code repository, downloadable weights, or hosted inference endpoint for VirtualCRISPR has been released.

#Applications

VirtualCRISPR targets the target-discovery stage of drug development, where functional genomics teams must decide which screen hits justify follow-up. By ranking hits on novelty and tractability, it helps prioritize candidates that are both mechanistically informative and druggable, reducing the time from a completed screen to a validated, actionable target. In its introducing study this shortened the path from screen to druggable hit and pointed directly to existing pharmacology: the ALOX5 inhibitor Zileuton and the OXTR antagonist Cligosiban, delivered topically, matched a systemic anti-IL17RA antibody in an imiquimod-induced psoriasis model. The approach generalizes to any pooled screen where separating novel biology from well-trodden pathways is the limiting step.

#Impact

VirtualCRISPR illustrates how a pretrained language model can be layered onto functional genomics to accelerate therapeutic discovery, turning a raw list of screen hits into a prioritized set of candidate drug targets that held up under molecular, organotypic, and animal-model validation. By surfacing ALOX5 and OXTR as psoriasis targets with little prior disease association and then confirming them experimentally, the work provides a concrete blueprint for integrating AI-driven target prioritization with large-scale genetic screens. Its current limitations are those of a recent preprint: the results await peer review, and no code, weights, or model documentation have been released, so independent reproduction of VirtualCRISPR itself is not yet possible. The demonstration also rests on a single disease context, and broader validation across screens and phenotypes will determine how generally the novelty-filtering approach transfers.

Citation

AI-Guided CRISPR Screen Accelerates Discovery of New Drug Targets

Zhao, C., et al. (2026) AI-Guided CRISPR Screen Accelerates Discovery of New Drug Targets. bioRxiv.

DOI: 10.64898/2026.02.26.708368

Recent citations

Papers that recently cited this model.

Not enough citation data yet.

Top citations

The most-cited papers that cite this model.

Not enough citation data yet.

Related models

Models with similar goals, methods, or subject matter.

  • OpenCRISPR-1

    Profluent

    AI-designed CRISPR-Cas9 gene editor generated by protein language models trained on 1.2 million CRISPR operons and shown to edit the human genome.

    Protein
  • TwinCell

    DeepLife

    Large causal cell model trained on cancer perturbation data that generalizes zero-shot to patient-derived cells for therapeutic target prioritization.

    Single-cell
  • scGenePT

    Chan Zuckerberg Initiative

    Single-cell perturbation prediction model that adds gene-level language embeddings from NCBI, UniProt, and Gene Ontology to scGPT representations.

    Single-cell
  • crisprSFM

    ETH Zurich

    CRISPR off-target prediction model that scores gRNA-DNA specificity from sequence, framing guide-target recognition as cross-modal retrieval.

    DNA & Gene
  • GEMGen

    Westlake University / Microsoft Research Asia

    Generative language model for phenotype-driven drug discovery, proposing small-molecule structures from up- and down-regulated gene signatures.

    Small moleculeSingle-cell
  • OwkinZero

    Owkin

    Biological reasoning language model that post-trains Qwen3 with reinforcement learning from verifiable rewards for drug-discovery tasks.

    Language model

Citations

Total Citations0
Influential0
References41

Fields of citing research

Not enough data

Openness

bio.rodeo opennessClosed · low usability and reproducibility
12Closed
Usability — can I run it?7
Reproducibility — can I retrain it?14
Model Openness Framework
Unclassified
Restrictive license on core components

Tags

crisprdrug_discoveryfoundation_modelfunctional_genomicszero_shot

Resources

Research Paper