bio.rodeo
ModelsOrganizationsLeaderboardAbout
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Small molecule foundation models
Small moleculeProtein

StructureSAFE

Purdue University

Structure-based drug design language model fusing protein structural and evolutionary encoders with SAFE fragment tokens for hit-to-lead generation.

Released: July 2026

Structure-based drug design seeks to generate molecules tailored to a specific protein binding site. Two families of generative models dominate the task, and each carries a characteristic weakness. Three-dimensional graph-based models condition explicitly on pocket geometry but, constrained by the scarcity of protein–ligand co-crystal training data, frequently propose chemically implausible or synthetically inaccessible structures. Chemical language models reliably emit valid, drug-like molecules but struggle to fold in three-dimensional structural context, and because they operate on SMILES strings they cannot naturally express the fragment-level edits that define lead optimization.

StructureSAFE, developed by researchers in the Borch Department of Medicinal Chemistry and Molecular Pharmacology at Purdue University and introduced in a 2026 bioRxiv preprint, targets this trade-off directly. It is a structure-aware chemical language model that pairs protein structural and evolutionary encoders with the SAFE (Sequential Attachment-based Fragment Embedding) molecular representation. SAFE re-expresses a molecule as an ordered sequence of fragment blocks, which keeps fragment-level operations easy to express while retaining the validity and fluency of a sequence model. Conditioning generation on encoded protein structure and evolution brings target awareness into the language-model backbone.

A single pretrained-then-finetuned checkpoint handles both de novo hit identification and a comprehensive suite of lead-optimization subtasks, and it generalizes to protein targets never seen during training. This unified, target-conditioned framing distinguishes StructureSAFE from pipelines that treat hit generation and lead optimization as separate problems.

#Key Features

  • Unified hit-ID and lead optimization: One framework spans de novo hit identification and a comprehensive suite of lead-optimization subtasks rather than requiring separate, task-specific models.
  • Structure and evolution conditioning: Protein structural and evolutionary encoders inject target-specific context, giving a language model the pocket awareness usually reserved for 3D graph models.
  • SAFE fragment representation: Encoding molecules as ordered fragment blocks makes fragment-level edits natural while preserving the chemical validity of sequence-based generation.
  • Generalizes to unseen targets: A rigorously constructed held-out test set confirms drug-like, synthetically accessible outputs with competitive predicted binding affinity for previously unseen proteins.
  • Pretrain–finetune scheme: Broad pretraining followed by task finetuning yields marked gains in chemical plausibility over graph-based models that lack pretraining.

#Technical Details

StructureSAFE is a chemical language model built on the SAFE representation and conditioned on protein structural and evolutionary encoders, trained under a two-stage pretraining-and-finetuning scheme. On the MolGenBench benchmark it reports state-of-the-art results across multiple metrics, with the most pronounced gains in chemical plausibility relative to graph-based baselines that lack pretraining. On a held-out test set the model produces drug-like, synthetically accessible molecules with competitive predicted binding affinities for previously unseen targets, across both hit identification and lead optimization settings. Four in silico case studies on therapeutically relevant targets show generated molecules recapitulating key binding interactions of known high-affinity ligands while proposing new interactions and exploring previously unexplored regions of chemical space.

#Applications

StructureSAFE is aimed at medicinal-chemistry workflows in both hit identification and lead optimization campaigns. Given a protein target and its structure, computational chemists and drug-discovery teams can use the model to propose candidate molecules, then refine promising scaffolds through fragment-level lead-optimization operations within the same framework. Because a single checkpoint generalizes to targets outside its training set, it can be applied to novel proteins without per-target retraining, providing a source of high-quality starting molecules to augment the early stages of a discovery program.

#Impact

StructureSAFE bridges two previously separate approaches to structure-based generation, showing that a chemical language model can absorb three-dimensional target context while retaining the chemical validity and synthetic accessibility that graph-based generators often sacrifice. By unifying hit identification and lead optimization in one target-conditioned model, it points toward more integrated computational medicinal-chemistry pipelines. The reported results are in silico, the work is a preprint awaiting peer review, and no code or trained weights accompany it; the preprint is released under a non-commercial, no-derivatives license.

Citation

StructureSAFE: A structure-aware chemical language model for unified hit identification and lead optimization

Yang, B., et al. (2026) StructureSAFE: A structure-aware chemical language model for unified hit identification and lead optimization. bioRxiv.

DOI: 10.64898/2026.06.28.735128

Recent citations

Papers that recently cited this model.

Not enough citation data yet.

Top citations

The most-cited papers that cite this model.

Not enough citation data yet.

Related models

Models with similar goals, methods, or subject matter.

  • UniLingo3DMol

    StoneWise

    Pretrained language model for 3D molecule generation in protein pockets, unifying de novo and fragment-based drug design in one multi-task framework.

    Small molecule
  • MolChord

    Beijing Zhongguancun Academy / University of Science and Technology of China

    Structure-based drug design model that generates ligands for a protein pocket, pairing a diffusion structure encoder with preference optimization.

    Small moleculeProtein
  • Molexar

    Peking University

    Multimodal molecular generation model for drug design, conditioned on properties, pharmacophores, protein sequences, or protein binding pockets.

    Small moleculeProtein
  • Sesame

    Tessel Biosciences

    Diffusion model that generates 3D small molecules conditioned on protein pockets and partial fragments encoded as continuous spatial density maps.

    Small moleculeProtein
  • SynPROTAC

    Sun Yat-sen University

    Designs synthesizable PROTAC degraders from reaction templates and purchasable building blocks, with reinforcement learning tuning the generator.

    Small molecule
  • Suiren-1.0

    Golab (SAIS Physics Lab)

    Molecular foundation models pretrained on density functional theory data, encoding 3D geometry and quantum behavior for ADMET and drug discovery.

    Small molecule

Citations

Total Citations0
Influential0
References16

Fields of citing research

Not enough data

Openness

bio.rodeo opennessClosed · low usability and reproducibility
10Closed
Usability — can I run it?7
Reproducibility — can I retrain it?14
Model Openness Framework
Unclassified
Restrictive license on core components

Tags

foundation_modelgenerativelead_optimizationtransformer

Resources

Research Paper