Manufacturing-aware generative sequence models whose parameters are DNA synthesis reaction conditions, so designs are made in vitro at petascale.
No providers recorded yet. Browse all providers
A generative antibody model can emit a quadrillion plausible CDRH3 loops overnight; a wet lab typically builds about 100,000 of them. The gap is chemistry, not compute: conventional synthesis writes each design out base by base as its own molecule, so cost scales with the number of designs. Degenerate-codon libraries escape that cost but are uniformly random, with no connection to any trained model.
Variational synthesis models close the gap by making the generative model and the synthesis protocol the same object. Each parameter of the model is an experimentally controlled quantity in a stochastic DNA oligosynthesis reaction — the mixture of bases delivered at a given coupling step, the reaction compartment a strand sits in. Training finds parameters θ* whose induced sequence distribution matches the data; running the corresponding reactions then draws samples from that distribution physically. Because every DNA molecule independently encounters a different series of reagents, each molecule is an independent sample, and the randomness inherent in the chemistry stands in for a random number generator. A trained model is therefore also a step-by-step manufacturing protocol.
JURA Bio introduced the idea theoretically in 2021 and implemented it at industrial scale with a commercial synthesis provider, posting the results in 2024 and publishing them in Nature Biotechnology in 2026, with collaborators at Harvard Medical School, the Wyss Institute and New York University. Three trained instances accompany the paper. The published models cover relatively short, mixed-length protein regions — CDRH3s of 4–40 amino acids, epitopes of 8–12, and a 40–60 residue polymerase segment; the authors attribute this to the oligosynthesis chemistry they chose rather than to variational synthesis itself, and do not present the approach as full-length protein generation.
The antibody instance was trained on 325 million unique human heavy-chain sequences drawn from 11,271 repertoires in the Observed Antibody Space database, excluding donors with autoimmune conditions. It reaches an average per-residue perplexity of 10.6 on held-out sequences, essentially matching its training perplexity. Against held-out data the in-silico BEAR Bayes factor is 1.31, compared with 1.02 for IgLM, and MMD is lower than IgLM's; 28% of in-silico samples clear the "clearly realistic" witness-function threshold, versus 49% for IgLM. Manufactured samples degrade only slightly — Bayes factor 1.39 and 39% clearly realistic from a 150 nmol, 9 × 10^16 molecule library — while an NNK degenerate-codon library of comparable diversity scores 0.08%. The epitope instance, trained on 2 million peptides screened by NetMHCpan-4.1 for HLA-A*02:01 binding, reaches perplexity 14.2 against a 11.4 floor (NNK: 20.6), and 36% of its manufactured sequences are predicted strong binders versus 0.7% for NNK. The polymerase instance, trained on 250,000 Taq C-terminal completions sampled from ProGen2, reaches perplexity 3.54 against a 2.57 floor. Roughly 10^16 realistic antibody designs cost about $10^3 to build, against an estimated $10^15 by per-design synthesis.
The output is a physical library rather than a downloadable checkpoint, which suits campaigns where the screen, not the design step, is the bottleneck. In the published work the antibody library was cloned into a second-generation scFv chimeric antigen receptor backbone, expressed in human cell lines and screened against multiplexed HLA-presented intracellular proteins, yielding candidate CARs against targets that are inaccessible to conventional surface-binder discovery. The epitope library supplies antigen candidates for T-cell vaccine work, and the polymerase library supports enzyme engineering for fidelity, thermostability or non-natural substrates.
Variational synthesis moves the constraint on model-driven protein design from what can be synthesized to what can be assayed — a reframing that matters most for methods, such as active learning over screens, that are starved by library sizes of 10^5. The practical caveat is access: only the general method code is public, the specific architectures were customized around proprietary platform details owned by the commercial synthesis provider, and no trained parameters have been released, so reproducing a library requires that partnership. The authors also note the compounding-error risk when a synthesis model is trained on another generative model's samples rather than on measured sequences.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.