Max Planck Institute for Heart and Lung Research
m6A RNA modification site prediction across the transcriptome, using a CNN-Transformer hybrid to surface unannotated N6-methyladenosine sites.
N6-methyladenosine (m6A) is the most abundant internal modification of messenger RNA, shaping splicing, nuclear export, stability, and translation across the transcriptome. Despite deep experimental mapping efforts, the catalog of confidently annotated m6A sites remains incomplete, and many positions that carry the modification in specific cellular contexts have never been assigned a functional role. M6AFormer addresses this gap by predicting m6A sites directly from RNA sequence and prioritizing unannotated candidates that share the sequence and contextual signatures of known functional sites.
The model was developed by the Gu laboratory at the Max Planck Institute for Heart and Lung Research (Zhixin Niu, Chang Liu, and Lei Gu) and released as an open-source preprint. Its central contribution is not only accurate site classification but the identification of previously unreported candidate sites, several of which the authors connect to disease-relevant genetic variation and GWAS loci. The team experimentally confirmed a previously undescribed m6A modification in NEU4 mRNA and demonstrated its functional consequences for cancer cell behavior, illustrating how the model's predictions translate into testable wet-lab hypotheses.
M6AFormer sits within a growing family of sequence-based epitranscriptomic predictors, distinguishing itself through a focus on discovery of functional, unannotated sites rather than reclassification of already-catalogued positions.
M6AFormer couples a convolutional feature extractor with a lightweight Transformer encoder that operates on windows of RNA sequence centered on candidate adenosines. Checkpoints are trained at 201 bp and 801 bp window sizes, and negatives are drawn either at random or restricted to the DRACH consensus motif to sharpen discrimination against sequence-matched decoys. Training data are organized into per-cell-line and pooled site sets, split with a 95:5 external partition for evaluation alongside an internal train/validation division used for early stopping. The published preprint reports improved site-prediction performance relative to existing m6A predictors and uses the trained models to scan the human transcriptome for candidate sites, which are then intersected with genetic-variation and GWAS resources to nominate functionally relevant modifications. The pretrained weights are bundled directly inside the installable package.
M6AFormer is aimed at epitranscriptomics researchers who need to scan transcripts or whole transcriptomes for candidate m6A sites without running new experiments. By prioritizing unannotated sites that resemble functional modifications and by connecting predictions to disease-associated variants and GWAS loci, the model serves as a hypothesis-generation tool for RNA biologists, cancer researchers, and human-genetics groups seeking to link epitranscriptomic marks to phenotype—as demonstrated by the follow-up validation of an m6A site in NEU4.
By foregrounding the discovery of unannotated functional sites, M6AFormer extends m6A prediction from confirmatory scoring toward active expansion of the epitranscriptome map, and its MIT-licensed release with API, CLI, and web interfaces lowers the barrier to routine transcriptome-wide scanning. The model is specialized for a single modification type and works from sequence alone, and its results are reported in a preprint that has not yet completed peer review, so predicted sites remain candidates for experimental confirmation of the kind the authors carried out for NEU4.
Niu, Z., et al. (2026) M6AFormer Prioritizes Unannotated Functional m6A Candidate Sites in the Human m6A Epitranscriptome. bioRxiv.
DOI: 10.64898/2026.07.10.737679Papers that recently cited this model.
The most-cited papers that cite this model.
Not enough data