Synthesis-aware generative model for small molecules that proposes drug-like analogs with forward routes from commercially available building blocks.
No providers recorded yet. Browse all providers
Ask a generative model for analogs of a lead compound and it will happily move a fluorine one position around a ring. That single edit can take a molecule from a one-step amide coupling to a route nobody has ever run — the structure passes every property filter and still never reaches a plate. Synthesis-aware generative models answer this by generating over reactions and purchasable building blocks rather than over atoms and bonds, so a route comes attached to each proposal by construction.
onepot C1 (Preview) is the first generative model from Onepot AI, a 1.2-billion-parameter conditional model for drug-like small molecules released in May 2026 alongside the company's Designer interface. It generates on top of CORE v1.1, Onepot's enumerated space of synthesizable compounds, and is conditioned at sampling time on a reference molecule, substructures to preserve, property windows, or any combination of the three. Onepot names SynFormer, PrexSyn and ReaSyn as the prior work it builds on; SynLlama attacks the same problem from the route side, serializing synthetic pathways over Enamine building blocks as text.
What distinguishes C1 is where its feasibility prior comes from. Onepot operates its own automated synthesis lab, so the model is trained on closed-loop execution data — records of which reactions actually ran on that platform, on which substrate combinations — rather than on retrosynthesis heuristics alone. Onepot has published no paper describing C1; the model is documented through the company's launch note and product pages, while the compound space it generates over is described in a separate preprint.
C1 is a 1.2-billion-parameter model trained on top of CORE v1.1. The first CORE release enumerates 3.4 billion compounds from 320,108 curated building blocks across seven medicinal-chemistry reactions — amide coupling (HATU/T3P), Suzuki–Miyaura coupling, Buchwald–Hartwig amination, CDI-mediated urea synthesis, TCDI-mediated thiourea synthesis, N-alkylation and O-alkylation — with roughly 73% of the space Ro5-compliant and only 4.7% overlap with Enamine REAL. Alongside that enumerated space, C1 trains on proprietary execution data from Onepot's platform, and a reinforcement learning pipeline follows.
Onepot has not stated C1's architecture family, its generation representation, or the size of the proprietary synthesis dataset. It reports RL training curves qualitatively and publishes no benchmark comparison against SynFormer, PrexSyn, ReaSyn or any other baseline, so there are no comparative numbers to quote. The shipped model is a fixed hosted checkpoint: conditioning is applied at inference, and customer-specific post-training — adapting the model to internal scoring models, property filters and preferred chemistries — is offered separately as a commercial service.
The intended uses are hit exploration, hit-to-lead and lead optimization: a medicinal chemist supplies a query structure and a property box, and receives makable neighbors ranked by feasibility rather than by structural plausibility alone. Because Onepot both designs and synthesizes, a selected compound can go from generation to a synthesis quote inside one system, with a median lead time under ten business days. Designer ships presets for common intents — lead optimization, drug-likeness, scaffold hopping — while the API path suits teams folding generation into an existing discovery pipeline.
C1's significance is less the model than the loop it closes: a generative model whose synthesizability prior is fit to data produced by the lab that will make its outputs, rather than to a static reaction template library. The chemical space underneath it already circulates beyond Onepot's own tooling — Talus Bioscience screened all 20,431 human proteins against the CORE library with Ptarmigan-1. The caveats are substantial. C1 is labelled a Preview, no weights or code are public, access is commercial, and no published evaluation of the model exists, so its feasibility and diversity claims rest entirely on the vendor's own reporting against its own compound space.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.