bio.rodeo
ModelsOrganizationsLeaderboardAbout
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
DNA & Gene foundation models
DNA & GeneSingle-cell

LEAF-1

McGill University / UCSF

Genomics foundation model that represents individual DNA fragments in a learned semantic space for cell-free DNA cancer detection and cell typing.

Released: July 2026

LEAF-1 is a foundation model for genomics that operates on individual DNA fragments rather than the reference-coordinate summaries most assays are reduced to. Sequencing assays such as ATAC-seq and cell-free DNA (cfDNA) profiling generate billions of short DNA fragments, but conventional pipelines aggregate those reads into coordinate-based features—binned coverage, peaks, or interval counts—discarding the biochemical detail carried by each molecule's exact sequence and cleavage boundaries. LEAF-1 was built to keep that information, taking a coordinate-free view in which the fragment itself is the unit of analysis.

The model, introduced in the preprint "Semantic fragment representations for coordinate-free analysis of genomics data," was developed by Hamed Najafabadi's group at McGill University with collaborators at the University of California, San Francisco, including Hani Goodarzi. It represents each DNA molecule as a point in a learned semantic space defined jointly by sequence context, the assay modality the fragment came from, and explicit cleavage-boundary tokens marking where the molecule was cut. This shared embedding space lets fragments from different experiments and assay types be compared and pooled directly.

LEAF-1 sits alongside a growing class of genomics and epigenomics foundation models, but is distinguished by its fragment-level, assay-agnostic representation and by the demonstration that its frozen embeddings transfer to new tasks and patient cohorts without retraining.

#Key Features

  • Fragment-level representation: Each DNA molecule is embedded individually, preserving per-fragment sequence and cut-site information that coordinate-based aggregation discards.
  • Cleavage-boundary tokens: The representation explicitly encodes where a fragment's ends fall, capturing the fragmentation patterns that reflect chromatin state and nuclease activity.
  • Assay-agnostic corpus: Pretraining spans bulk ATAC-seq, single-cell ATAC-seq, and cell-free DNA, so a single embedding space serves accessibility profiling and liquid-biopsy data alike.
  • Zero-shot transfer: A classifier trained on the frozen model generalizes to an unseen cancer type, detecting clear cell renal cell carcinoma from plasma without any retraining.
  • Sparse-data cell typing: Human cell types are classified from as few as roughly 1,000 fragments per cell using mean-pooled fragment embeddings.

#Technical Details

LEAF-1 was pretrained on approximately 58 billion DNA fragments drawn from bulk ATAC-seq, single-cell ATAC-seq, and cell-free DNA profiles. Every fragment is mapped into a common semantic space conditioned on its sequence context, its assay of origin, and explicit cleavage-boundary tokens; cell- or sample-level representations are obtained by mean-pooling the embeddings of the fragments belonging to that cell or sample. On downstream evaluations, LEAF-1 classifies human cell types from sparse single-cell data using around 1,000 fragments per cell, reaches an area under the ROC curve of 0.95 for cell-free DNA cancer detection, and—applying a frozen classifier to an independent cohort—achieves an AUC of 0.83 in detecting clear cell renal cell carcinoma, a cancer type withheld from that classifier. These results show that fragment-level modeling retains biochemical and disease-associated signal that coordinate-based aggregation loses.

#Applications

LEAF-1 targets researchers in cancer genomics, liquid biopsy, and single-cell epigenomics. Its cfDNA results point toward non-invasive cancer detection and classification from blood plasma, including generalization to cancer types not seen during classifier training, while its cell-typing results support annotation of sparse single-cell ATAC-seq experiments where per-cell fragment counts are low. Because the model emits reusable fragment and sample embeddings, it can serve as a frozen feature extractor that downstream classifiers plug into without task-specific pretraining.

#Impact

By modeling DNA one fragment at a time, LEAF-1 reframes how heterogeneous genomics assays can be represented in a single learned space and demonstrates that coordinate-free embeddings transfer across tasks, assay types, and patient cohorts. The frozen-checkpoint result on a withheld cancer type is notable evidence that the representation captures generalizable disease signal rather than dataset-specific artifacts. Several caveats apply at the time of writing: LEAF-1 is a preprint awaiting peer review, its reported benchmarks are computational, the paper is released under a non-commercial no-derivatives license (CC BY-NC-ND), and no public code or pretrained weights accompany the preprint, so independent reproduction and reuse are not yet possible.

Citation

Semantic fragment representations for coordinate-free analysis of genomics data

Heydari, H., et al. (2026) Semantic fragment representations for coordinate-free analysis of genomics data. bioRxiv.

DOI: 10.64898/2026.07.09.737627

Recent citations

Papers that recently cited this model.

Not enough citation data yet.

Top citations

The most-cited papers that cite this model.

Not enough citation data yet.

Related models

Models with similar goals, methods, or subject matter.

  • cfRNA-ICL

    Eigen Bio

    In-context learning model for cell-free RNA, meta-trained on synthetic tasks from a cfRNA structural causal model for few-shot cancer classification.

    Single-cellRNA
  • GenBloom

    Helmholtz Munich / LMU Munich

    Genetically aligned foundation model for blood smear cytology that links single-cell morphology to the chromosomal aberrations behind AML and APL.

    Pathology
  • RNABag

    HomiGen Intelligence Technology Co., Ltd.

    Transcriptome foundation model for precision oncology, generalizing zero-shot across tissue, plasma cfRNA, and tumor-educated platelet modalities.

    Single-cell
  • TESSERA

    Weill Cornell Medicine

    Self-supervised foundation model that embeds cancer genomes from somatic SNVs and copy-number alterations across 33 tumor types for tumor subtyping.

    DNA & Gene
  • BioMed Multi-Omic

    IBM Research

    Open-source framework for building RNA and DNA foundation models, featuring WCED pretraining for transcriptomics and SNP-aware encoding for genomics.

    DNA & Gene

Citations

Total Citations0
Influential0
References32

Fields of citing research

Not enough data

Openness

bio.rodeo opennessClosed · low usability and reproducibility
4Closed
Usability — can I run it?7
Reproducibility — can I retrain it?0
not reproducible
Model Openness Framework
Unclassified
Restrictive license on core components

Tags

cancer_detectioncell_type_annotationcfdnachromatinfoundation_modelrepresentation_learningself_supervisedzero_shot

Resources

Research Paper