Genomics foundation model that represents individual DNA fragments in a learned semantic space for cell-free DNA cancer detection and cell typing.
LEAF-1 is a foundation model for genomics that operates on individual DNA fragments rather than the reference-coordinate summaries most assays are reduced to. Sequencing assays such as ATAC-seq and cell-free DNA (cfDNA) profiling generate billions of short DNA fragments, but conventional pipelines aggregate those reads into coordinate-based features—binned coverage, peaks, or interval counts—discarding the biochemical detail carried by each molecule's exact sequence and cleavage boundaries. LEAF-1 was built to keep that information, taking a coordinate-free view in which the fragment itself is the unit of analysis.
The model, introduced in the preprint "Semantic fragment representations for coordinate-free analysis of genomics data," was developed by Hamed Najafabadi's group at McGill University with collaborators at the University of California, San Francisco, including Hani Goodarzi. It represents each DNA molecule as a point in a learned semantic space defined jointly by sequence context, the assay modality the fragment came from, and explicit cleavage-boundary tokens marking where the molecule was cut. This shared embedding space lets fragments from different experiments and assay types be compared and pooled directly.
LEAF-1 sits alongside a growing class of genomics and epigenomics foundation models, but is distinguished by its fragment-level, assay-agnostic representation and by the demonstration that its frozen embeddings transfer to new tasks and patient cohorts without retraining.
LEAF-1 was pretrained on approximately 58 billion DNA fragments drawn from bulk ATAC-seq, single-cell ATAC-seq, and cell-free DNA profiles. Every fragment is mapped into a common semantic space conditioned on its sequence context, its assay of origin, and explicit cleavage-boundary tokens; cell- or sample-level representations are obtained by mean-pooling the embeddings of the fragments belonging to that cell or sample. On downstream evaluations, LEAF-1 classifies human cell types from sparse single-cell data using around 1,000 fragments per cell, reaches an area under the ROC curve of 0.95 for cell-free DNA cancer detection, and—applying a frozen classifier to an independent cohort—achieves an AUC of 0.83 in detecting clear cell renal cell carcinoma, a cancer type withheld from that classifier. These results show that fragment-level modeling retains biochemical and disease-associated signal that coordinate-based aggregation loses.
LEAF-1 targets researchers in cancer genomics, liquid biopsy, and single-cell epigenomics. Its cfDNA results point toward non-invasive cancer detection and classification from blood plasma, including generalization to cancer types not seen during classifier training, while its cell-typing results support annotation of sparse single-cell ATAC-seq experiments where per-cell fragment counts are low. Because the model emits reusable fragment and sample embeddings, it can serve as a frozen feature extractor that downstream classifiers plug into without task-specific pretraining.
By modeling DNA one fragment at a time, LEAF-1 reframes how heterogeneous genomics assays can be represented in a single learned space and demonstrates that coordinate-free embeddings transfer across tasks, assay types, and patient cohorts. The frozen-checkpoint result on a withheld cancer type is notable evidence that the representation captures generalizable disease signal rather than dataset-specific artifacts. Several caveats apply at the time of writing: LEAF-1 is a preprint awaiting peer review, its reported benchmarks are computational, the paper is released under a non-commercial no-derivatives license (CC BY-NC-ND), and no public code or pretrained weights accompany the preprint, so independent reproduction and reuse are not yet possible.
Heydari, H., et al. (2026) Semantic fragment representations for coordinate-free analysis of genomics data. bioRxiv.
DOI: 10.64898/2026.07.09.737627Papers that recently cited this model.
The most-cited papers that cite this model.
Not enough data