Single-cell multiomic foundation model that transfers pan-cancer RNA-ATAC regulatory structure into RNA-only tumour datasets via low-rank adapters.
No providers recorded yet. Browse all providers
Single-cell assays that measure RNA and chromatin accessibility in the same cell link transcriptional state directly to regulatory-element usage. Matched RNA–ATAC datasets, however, remain far scarcer than scRNA-seq atlases, and the gap is most costly in tumours, where regulatory signals are weak, sparse and often confined to rare malignant states or local microenvironmental niches. Expression-only analysis misses those relationships; unconstrained recovery converts technical variation into plausible-looking networks.
IMAS — the Integrative Multiomic Augmentation System — learns regulatory structure once from a pan-cancer corpus of matched single-cell RNA and ATAC profiles, then transfers it into target datasets that carry only RNA. It was developed by Wu Deyang, Takashi Yamashiro and Toshihiro Inubushi in the Department of Orthodontics and Dentofacial Orthopedics at Osaka University, and posted to bioRxiv in April 2026.
The output is not a completed expression matrix. IMAS emits multi-layer target dependencies (MLTDs): compact, target-specific regulatory architectures that retain a candidate signal only when it stays supported across matched multiomic priors, adapted RNA–transcription-factor–regulatory-element relations, cell–cell communication context and perturbation-sensitive local structure. This places it alongside single-cell foundation models such as scGPT and perturbation-response models such as GEARS and CPA, against which it is benchmarked, while targeting a different deliverable — a prioritized, traceable hypothesis set rather than a predicted expression profile.
A multi-relation prior graph anchors the model: TF–regulatory-element edges come from FIMO motif scanning of merged hg38 peaks, regulatory elements link to up to three nearest genes within 200 kb with distance-decayed weights, and composed TF–gene edges keep the top 2,000 targets per transcription factor. The backbone is trained in a leave-one-dataset-out framework over paired cells, with per-cell token budgets of 2,048 RNA features and 1,024 regulatory-element features, combining cross-modal alignment with prior-graph link supervision under AdamW with gradient clipping. The pan-cancer corpus is broad but uneven — medulloblastoma, prostate cancer and glioma contribute the most matched cells, and the ten largest cohorts account for 67.7% of cells.
Transfer freezes that backbone and partitions target cells by spherical k-means, then trains adapters of rank 32 against three objectives — RNA-to-TF, RNA-to-RE and RNA-to-RNA reconstruction — with TF and RE margins of 0.2, AdamW for 10,000 steps at a learning rate of 7 × 10⁻⁴, and checkpoint selection on validation TF AUC. Faithfulness is assessed by progressive deletion and retention of top-ranked support genes against weight- and degree-matched null selections. On recovery of perturbation-associated expression direction, IMAS exceeded CPA, scGPT and GEARS, with same-trend recovery rising as target-gene connectivity increased.
IMAS is aimed at cancer biologists holding RNA-only single-cell data who need a short, testable list of regulatory mechanisms. In colorectal cancer it resolved a SOD2-associated architecture distinct from native expression and conventional co-expression, and a LAMB1-centred analysis restored an extracellular-matrix programme weakly represented in the raw matrix. In head and neck squamous cell carcinoma, SOX2-centred dependencies separated malignant cells into basal-like, MYC-associated and RTK–MAPK stress-response programmes. In renal cancer, the model nominated a tumour–vascular coupling that Xenium in situ profiling of clear-cell carcinoma tissue supported through endothelial regions with mesenchymal and vascular-abnormality features. Glioma and osteosarcoma served as further transfer targets.
IMAS reframes tumour multiomic augmentation as recovery of compact regulatory architectures rather than matrix completion, and its ablations are a useful data point for the field: matched cross-layer supervision, not raw cell count, carried the transferable signal. Adoption is untested. The work is an unreviewed preprint from a single group with no third-party replication, code and analysis scripts are promised on GitHub only upon publication and no repository or pretrained checkpoint has been released, and the eight-patient renal cohort awaits a GEO accession. The authors are explicit that virtual perturbation prioritizes hypotheses rather than substituting for genetic or pharmacological intervention, and that spatial correspondence supports tissue-level plausibility without proving individual regulatory mechanisms.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.