Aligns structure, binding-pocket, text and molecular-dynamics encoders to a protein sequence anchor, giving frozen embeddings that transfer widely.
No providers recorded yet. Browse all providers
A sequence encoder never sees the pocket geometry that decides whether a ligand binds, and a structure encoder never sees the UniProt annotation naming the reaction the enzyme catalyses. Training one joint model on all of it usually demands proteins where every modality is present at once, and those slices barely overlap — so the requirement throws most of the data away.
OneProt sidesteps that by borrowing the ImageBind recipe. Every modality gets its own encoder, and each is aligned by an InfoNCE contrastive loss to a single anchor — the amino acid sequence — using only pairs, never complete tuples. A protein with a structure but no detected pocket still trains the sequence-structure alignment.
The model was built at the Jülich Supercomputing Centre of Forschungszentrum Jülich with Helmholtz AI, Helmholtz Munich, the Technical University of Munich and Heinrich Heine University Düsseldorf, appearing as a preprint in November 2024 and in PLOS Computational Biology a year later. Two checkpoints are released: OneProt-5, spanning all five modalities, and OneProt-4, which omits the structure-token encoder. A follow-up study from the same group adds all-atom, time-resolved molecular dynamics trajectories as a sixth modality, re-running the ablation grid with and without them.
The sequence anchor is a frozen ESM-2 650M encoder. Around it sit a 35M-parameter ESM2 transformer trained from scratch on Foldseek structure tokens, two 2.6M all-atom ProNet graph networks for backbone structure and for pockets, and a BiomedBERT text encoder adapted with LoRA — 802.2M parameters in total, of which 40.6M are trainable. Pretraining ran for 33,000 optimizer steps on 64 NVIDIA A100 GPUs of the JUWELS Booster supercomputer, over OpenProteinSet and UniProtKB/Swiss-Prot clustered at 50% sequence identity: 1.04M sequences, 1M structure tokens, 656K structure graphs, 546K text annotations and 341K pockets. Across ten downstream benchmarks over six runs each, OneProt-4 reaches 0.668 Spearman on FLIP thermostability, 88.8% accuracy on HumanPPI and 0.877 Fmax on Enzyme Commission prediction, matching or beating SaProt, ESM-3, OpenFold in SoloSeq mode and ProTrek — which trained on 40M paired datapoints against OneProt's 1M.
The MD modality reuses the latent encoder of MDGen — a Scalable Interpolant Transformer whose timewise attention is replaced by long-context Hyena layers — loaded with published weights and held frozen, its hidden states pooled over frames, then residues, and projected into the shared 1024-dimensional space. Trajectories from mdCATH, GPCRmd and ATLAS, resampled to 1 ns spacing, cover 5,005 unique sequences held out of the downstream evaluation sets. Pretraining for 32,900 steps on 128 A100 GPUs, updating only the projection weights, raises thermostability for every variant — OneProt-4 from 0.668 to 0.684 Spearman — and helps most on subcellular localisation and metal-ion binding where graph structure is absent, while Enzyme Commission and GO molecular function are unchanged.
The practical output is an embedding service: a sequence goes in, a vector comes out already carrying structural, pocket and functional-annotation signal, ready for a small classifier. That suits enzyme function assignment, subcellular localisation, thermostability ranking and metal-ion binding screens where a lab has a few thousand labelled proteins and no budget to fine-tune a large encoder — with the MD-augmented variants preferred when the property turns on flexibility. Cross-modal retrieval supports the inverse query: finding proteins whose pockets match a described function, the shape of a hit-list problem in drug repurposing.
OneProt demonstrates that modality coverage can substitute for scale: a 40.6M trainable-parameter alignment over about a million pairs holds its own against contrastive models trained on forty times the data. Weights for both checkpoints are on HuggingFace under an MIT license — OneProt-4 downloadable outright, OneProt-5 behind a click-through agreement — with pretraining splits on Zenodo, though inference means cloning the repository rather than installing a library. A further Zenodo record holds the downstream checkpoints and benchmark splits, but no MD-extended pretrained checkpoint is actually published. The MD results mark the limit of piling on modalities: with trajectories for only 5,005 proteins, dynamics fills gaps left by missing structure but adds nothing once all five static modalities are present.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.