bio.rodeo
ModelsOrganizationsProvidersLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsProvidersLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
models / protein / dodock
ProteinSmall molecule
Deep OriginReleased August 2026

DODock

Hybrid diffusion-and-physics docking model that predicts protein-ligand binding poses and generalizes out-of-distribution for virtual screening.

4Openness

Where to run it

No providers recorded yet. Browse all providers

DODockProteinDeep Origin

Virtual screening promises access to tens of billions of synthesizable compounds, yet it is rarely used as the primary route to new chemical matter. The authors trace this gap to a persistent tradeoff in the underlying methods: classical docking such as AutoDock Vina generalizes across targets but is capped in accuracy by simple scoring functions, while recent machine-learning docking is highly expressive yet generalizes poorly to novel molecules and pockets, with reported accuracy often inflated by train-test leakage.

DODock, from Deep Origin, attacks this tradeoff with a hybrid design that pairs a learned diffusion pose sampler with physics-based refinement. Introduced in an August 2026 bioRxiv preprint alongside its companion scoring model DOScore — DODock predicts binding poses, DOScore ranks binding affinity — the model is trained once and applied zero-shot to unseen protein families and ligand chemistries rather than refit per target. Strict protein- and ligand-similarity train-test splits are used throughout to demonstrate out-of-distribution generalization, directly rebutting the leakage critique the authors level at prior ML docking systems such as AlphaFold 3, Boltz-1, and Chai-1.

The headline evidence is prospective: a blind DODock prediction of a drug candidate bound to PCSK9 recovered the pose to 1.2 Å RMSD against a crystal structure solved only afterward.

#Key Features

  • Diffusion sampler plus physics refinement: A learned diffusion model proposes candidate poses that are refined against an adapted AutoDock Vina-style potential using an enhanced scatter-search global optimizer, combining ML expressiveness with a physically grounded energy function.
  • Parsimonious, generalizable scoring: The physics potential carries only around 80 free parameters, a deliberately small footprint that resists overfitting and helps a single checkpoint transfer across chemically distinct targets.
  • Learned pose ranking: A dedicated ranking network selects among refined poses, replacing the hand-tuned heuristics of classical docking.
  • Single zero-shot checkpoint: One trained model is evaluated without per-target retraining across CASF-2016, Runs N' Poses, OpenBind, and PoseBusters, and deployed unchanged in prospective campaigns.
  • Leakage-controlled evaluation: Benchmarks use strict protein- and ligand-similarity splits and out-of-distribution holdouts, so reported accuracy reflects generalization rather than memorized training neighbors.

#Technical Details

DODock is a hybrid ML/physics docking framework. A diffusion network samples ligand poses in the protein pocket; these are refined by minimizing an adapted Vina-style scoring function through an enhanced scatter-search procedure, and a learned network ranks the resulting poses. The scoring potential's roughly 80 free parameters keep the model deliberately low-capacity relative to fully learned docking, which the authors argue is central to its generalization. Evaluation spans redocking and pose-prediction benchmarks — CASF-2016, Runs N' Poses, OpenBind (including an Enterovirus A71 2A protease case chosen to be dissimilar from the training data), and PoseBusters stratified by protein family. Prospective validation includes the blind PCSK9 pose at 1.2 Å RMSD and four virtual-screening campaigns against an ectoenzyme (CD73), a kinase (IRAK4), an extended-substrate protease (FXI), and an allosteric protein-protein interface (IL17). On CD73, a historically difficult screening target, the campaign reported roughly a hundredfold improvement in hit rate over a recent machine-learning screen.

#Applications

DODock targets structure-based virtual screening and pose prediction for hit discovery and lead optimization, where a pocket is known and the task is to predict which candidate molecules bind and in what geometry. Its four prospective campaigns yielded chemically novel, biochemically and cellularly active inhibitors across enzyme, protease, and protein-protein-interface targets, and the blind PCSK9 result demonstrates reliability on a genuinely unseen system. Paired with DOScore for affinity ranking, it is intended as a screening engine for interrogating ultralarge synthesizable libraries.

#Impact

DODock reframes the long-standing plateau in virtual-screening accuracy as a consequence of the accuracy-generalization tradeoff rather than a fundamental ceiling, and its leakage-controlled, prospectively validated results are a pointed response to concerns that ML docking benchmarks overstate real-world performance. The work is a single-company technical report that has not completed peer review, and the reported hit-rate gains await independent replication. No public code, weights, or hosted endpoint has been released; Deep Origin exposes docking through commercial products including Balto, DO Studio, and the Deep Origin API, so the model is documented in the preprint but not independently runnable.

At a glance

Released
August 2026
Category
Protein
Organization
Deep Origin

Links

bioRxiv Preprint

Tags

diffusiongenerativemolecular_dockingvirtual_screeningzero_shot

Something wrong?

Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.