Structure-based drug design language model fusing protein structural and evolutionary encoders with SAFE fragment tokens for hit-to-lead generation.
Structure-based drug design seeks to generate molecules tailored to a specific protein binding site. Two families of generative models dominate the task, and each carries a characteristic weakness. Three-dimensional graph-based models condition explicitly on pocket geometry but, constrained by the scarcity of protein–ligand co-crystal training data, frequently propose chemically implausible or synthetically inaccessible structures. Chemical language models reliably emit valid, drug-like molecules but struggle to fold in three-dimensional structural context, and because they operate on SMILES strings they cannot naturally express the fragment-level edits that define lead optimization.
StructureSAFE, developed by researchers in the Borch Department of Medicinal Chemistry and Molecular Pharmacology at Purdue University and introduced in a 2026 bioRxiv preprint, targets this trade-off directly. It is a structure-aware chemical language model that pairs protein structural and evolutionary encoders with the SAFE (Sequential Attachment-based Fragment Embedding) molecular representation. SAFE re-expresses a molecule as an ordered sequence of fragment blocks, which keeps fragment-level operations easy to express while retaining the validity and fluency of a sequence model. Conditioning generation on encoded protein structure and evolution brings target awareness into the language-model backbone.
A single pretrained-then-finetuned checkpoint handles both de novo hit identification and a comprehensive suite of lead-optimization subtasks, and it generalizes to protein targets never seen during training. This unified, target-conditioned framing distinguishes StructureSAFE from pipelines that treat hit generation and lead optimization as separate problems.
StructureSAFE is a chemical language model built on the SAFE representation and conditioned on protein structural and evolutionary encoders, trained under a two-stage pretraining-and-finetuning scheme. On the MolGenBench benchmark it reports state-of-the-art results across multiple metrics, with the most pronounced gains in chemical plausibility relative to graph-based baselines that lack pretraining. On a held-out test set the model produces drug-like, synthetically accessible molecules with competitive predicted binding affinities for previously unseen targets, across both hit identification and lead optimization settings. Four in silico case studies on therapeutically relevant targets show generated molecules recapitulating key binding interactions of known high-affinity ligands while proposing new interactions and exploring previously unexplored regions of chemical space.
StructureSAFE is aimed at medicinal-chemistry workflows in both hit identification and lead optimization campaigns. Given a protein target and its structure, computational chemists and drug-discovery teams can use the model to propose candidate molecules, then refine promising scaffolds through fragment-level lead-optimization operations within the same framework. Because a single checkpoint generalizes to targets outside its training set, it can be applied to novel proteins without per-target retraining, providing a source of high-quality starting molecules to augment the early stages of a discovery program.
StructureSAFE bridges two previously separate approaches to structure-based generation, showing that a chemical language model can absorb three-dimensional target context while retaining the chemical validity and synthetic accessibility that graph-based generators often sacrifice. By unifying hit identification and lead optimization in one target-conditioned model, it points toward more integrated computational medicinal-chemistry pipelines. The reported results are in silico, the work is a preprint awaiting peer review, and no code or trained weights accompany it; the preprint is released under a non-commercial, no-derivatives license.
Yang, B., et al. (2026) StructureSAFE: A structure-aware chemical language model for unified hit identification and lead optimization. bioRxiv.
DOI: 10.64898/2026.06.28.735128Papers that recently cited this model.
The most-cited papers that cite this model.
Not enough data