Antisense oligonucleotide activity prediction from sequence, position-specific chemistry, dose, and cell context. Spearman 0.5970 on ASO Atlas.
No providers recorded yet. Browse all providers
Antisense oligonucleotides (ASOs) are short, chemically modified single-stranded nucleic acids that hybridize to a target transcript and recruit RNase H to degrade it, reaching disease genes that small molecules and antibodies cannot. Choosing which candidate to synthesize is the bottleneck, because activity is not set by complementarity alone: it depends on the pattern of sugar and backbone modifications along the oligonucleotide, the accessibility of the binding site within its flanking RNA, the dose and transfection protocol, and the transcriptional state of the assayed cell line. Most screening tools capture only a subset, so predictions weaken on a new gene, cell line, or modification pattern.
ASOCompass, developed at Shanghai Jiao Tong University with a collaborator at the University of Toronto and posted as a preprint in August 2026, models these determinants jointly in one activity-prediction network. It pairs nucleotide identity with modified-monomer chemistry position by position, encodes the target RNA with its flanking sequence, attends from each ASO position to the binding site, and conditions the output on dose, transfection method, target gene, and cell line. Auxiliary regression on monomer physicochemical properties and ASO–target thermodynamics injects priors that sparse inhibition labels cannot supply.
The design goal is transferability rather than in-distribution accuracy: evaluation spans four distribution shifts and, more tellingly, two target genes withheld entirely from training, where only the prediction head is refit.
ASO and target-RNA sequences are encoded by RiNALMo with shared weights, and a gated residual fusion module aligns the resulting nucleotide embeddings with the Uni-Mol2 chemistry embeddings position by position. Cell context comes from genome-wide DepMap expression profiles passed through BulkFormer, yielding a mean-pooled cell-state vector and the target gene's own representation, both adapted by learnable prototype banks. Four streams — ASO, full target, binding site, and interaction — are summarized by multi-query attention pooling and passed with the context features to an MLP predicting inhibition percentage.
Training uses ASO Atlas, a patent-derived corpus of 188,521 RNase H gapmer records from 417 USPTO patents covering 334 target genes and 31 cell lines; filtering leaves 153,174 records, of which 28,726 are held out and partitioned by whether the drug, gene, and cell line were seen in training. ASOCompass reaches an overall Spearman correlation of 0.5970 against 0.5549 for OligoAI, the strongest ASO-specific baseline, and leads on all four shifts: 0.7207 on unseen drugs, 0.3609 on unseen genes, 0.5702 on unseen cell lines, and 0.5536 when both are unseen. Ablations attribute 0.0294 of that gain to biological context and 0.0200 to the chemistry encoder. On SOD1 and KLKB1, excluded from base training, head-only fine-tuning on 1,024 labels reaches 0.830 and 0.696.
ASOCompass is aimed at oligonucleotide discovery teams choosing which gapmers to order for a new target. Given a candidate walk across a transcript, it ranks designs under the dose, delivery route, and cell line planned for the assay, and separates chemistry patterns sharing a nucleotide sequence: chemical-property supervision raises pairwise accuracy on modification-only comparisons from 0.562 to 0.725 and cuts selection regret from 14.36 to 8.59 percentage points. For a gene with no screening history, one round of measurements refits the prediction head and reprioritizes the rest.
ASOCompass shows that ASO activity prediction improves when the assay itself is treated as input rather than noise, and that pretrained RNA, molecular, and transcriptomic encoders compose into one screening model. Auxiliary supervision earns its place: thermodynamic targets help most under gene and cell-line shift, chemical-property targets sharpen modification-level ranking. The work is a preprint awaiting peer review. Training and evaluation are retrospective and drawn from patent-reported inhibition percentages, heterogeneous in protocol and biased toward compounds worth patenting; accuracy on unseen genes stays modest; and the scope is RNase H gapmers, not splice-switching or steric-blocking ASOs. A code repository under the authors' MAGIC-AI4Med GitHub organization is referenced but not publicly accessible, and no trained weights have been released.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.