Generative chemistry ensemble fine-tuned on 2.4M RNA-small molecule binding measurements, designing analogues around a seed compound.
No providers recorded yet. Browse all providers
The public corpora that generative chemistry models learn from — ChEMBL, ZINC 250k, GEOM-Drugs — are the accumulated output of decades of protein-targeting medicinal chemistry. Seed one of those models with risdiplam, the only FDA-approved small molecule with a known RNA-modulating mechanism, and it returns analogues from the chemical space it knows, which is not the space that binds structured RNA. The Serna Bio GenAI platform changes the corpus rather than the architecture: SMILES language models fine-tuned on experimentally measured RNA–small molecule binding, so the chemistry proposed is already RNA-shaped.
The platform is an ensemble rather than a single network. Two SMILES BERT masked language models and a matched-molecular-pairs (MMP) transformation database were trained on an in-house corpus of RNA binders, and an unmodified pretrained SMILES BERT generator joins them to form what the authors call the Serna Bio GenAI platform ensemble. Generation is entirely ligand-dependent — a seed compound as a SMILES string is the only input, with no RNA structure required — and a predictive-model ensemble plus a multiparametric optimization (MPO) function then rank the output for synthesis. The work comes from Serna Bio with CompChem Solutions and the University of Michigan Biointerfaces Institute, and was first posted in May 2026. Serna Bio markets the stack commercially as Polaris.
The scope is deliberately narrow. The generators were selected to explore chemical space immediately around a provided seed — the move a medicinal chemist makes in lead optimization — and explicitly not to scaffold hop; because training and test sets were split randomly within RNA-binding chemical space, the authors note the models are less suitable for out-of-distribution generation.
The Serna Bio RNA–small molecule dataset holds more than 2.4 million binding datapoints collected across seven screening exercises using the Automated Ligand Identification System (ALIS) and SAMDI-ASMS, spanning 181,632 unique compounds against 85 RNA targets — G-quadruplexes, multi-way junctions and stem-loops within genes known to form secondary structure. Of the compounds screened, 5,726 (3.2%) bind at least one target. SMILES were canonicalized with RDKit; the BERT models were fine-tuned for 5 and 10 epochs at a learning rate of 0.0001, mask probability 0.15, batch size 8 and a random 20% held-out test fraction, with five known RNA binders withheld entirely as validation seeds. Given risdiplam as a seed, the ensemble produced 61 compounds with a mean Tanimoto similarity of 0.77 to the seed and 83.6% above 0.7, against 0% for both REINVENT (11,481 compounds, mean 0.15) and MolMIM (46 compounds, mean 0.18); 91.7% of its designs scored above 50 in a GOLD docking model of the SMN2–U1 duplex, versus 11.5% and 4.3%. The three design sets shared no compounds.
The platform is built for hit-to-lead and lead optimization in RNA-targeted drug discovery, where a team already has a binding hit and needs analogues that survive SAR cliffs. In a GLUT1 deficiency syndrome program, the ensemble generated 411 unique valid analogues from a hit series while a human medicinal chemist independently designed 26, with only 2 in common. Of 22 synthesized machine designs, 6 (27%) reached an EC25 below 1 µM for GLUT1 protein upregulation versus 2 of 17 (12%) human designs — a roughly threefold potency gain in one design-make-test cycle. The most potent design was confirmed by ALIS to bind the GLUT1 3'UTR motif, and raised both GLUT1 mRNA and cellular glucose uptake.
The result worth carrying forward is not the leaderboard margin but the demonstration that training-set modality, not architecture, was the binding constraint: an off-the-shelf SMILES BERT re-pointed at RNA-binding data designs chemistry that protein-trained generators do not reach. The caveats are substantial and the authors state them: the comparison ran REINVENT and MolMIM at default settings, docking is a proxy for activity, the experimental evidence is one program with a small sample, and the downstream predictive models and MPO components are described but not evaluated. The platform is closed — no code, weights or training data have been released, and the preprint has not been peer reviewed — so the approach is reproducible only by groups holding comparable RNA-binding screening data.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.