bio.rodeo
ModelsOrganizationsLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
Protein foundation models
Protein

ProTrek

Westlake University

Tri-modal protein language model aligning sequence, structure, and text in one embedding space for natural-language search over billions of proteins.

Released: June 2024

ProTrek is a tri-modal protein language model developed at Westlake University that jointly learns from three complementary data types: amino acid sequences, 3D structures, and natural-language functional descriptions. Published in Nature Biotechnology in 2025, the model addresses a fundamental gap in protein informatics — the inability to search the protein universe using all three modalities simultaneously. Traditional tools like MMseqs2 or Foldseek operate within a single modality; ProTrek bridges all three through a unified contrastive learning framework.

The architecture aligns sequence, structure, and function representations in a shared embedding space via three pairwise contrastive objectives: sequence-structure, sequence-function, and structure-function alignment. This design enables nine distinct cross-modal search tasks — every pairwise combination of the three modalities, in both directions, plus within-modality retrieval. A researcher can, for example, input a natural-language query such as "serine protease involved in blood coagulation" and retrieve structurally and functionally relevant proteins from a database of billions in seconds.

ProTrek is available in two size variants and ships with precomputed embeddings covering over five billion proteins from NCBI, GOPC, and OMG/MGnify metagenomic repositories, accessible through a public web server.

#Key Features

  • Tri-modal contrastive learning: Aligns sequence, structure, and function encoders in a single shared embedding space, supporting all nine pairwise cross-modal retrieval tasks without task-specific fine-tuning.
  • Natural-language protein search: Accepts plain-English functional queries against a precomputed index of 5+ billion proteins, lowering the barrier for biologists without computational expertise.
  • Retrieval speed and accuracy gains: Achieves 30x improvement in sequence-to-function retrieval and 60x in function-to-sequence retrieval relative to baseline methods, with overall search speed 100x faster than Foldseek or MMseqs2.
  • Strong transfer representations: Outperforms ESM-2 (650M parameters) on 9 of 11 supervised downstream prediction benchmarks, including binding site prediction, subcellular localization, and thermostability.
  • Billion-scale deployment: Precomputed FAISS-indexed embeddings for over five billion proteins are served through the ProTrek search server; local deployment against custom databases is also supported.

#Technical Details

ProTrek comprises three modality-specific encoders. The sequence encoder is based on the ESM-2 architecture (35M or 650M parameters depending on model variant). The structure encoder follows the Foldseek architecture (35M or 150M parameters). The text encoder is derived from BiomedNLP-PubMedBERT (130M parameters), providing semantic grounding in biomedical language. Total parameter counts are approximately 200M for ProTrek_35M and 930M for ProTrek_650M.

Training used Swiss-Prot (manually curated, high-quality annotations) and TrEMBL50 (large-scale automatic annotations) as primary datasets. Contrastive alignment is implemented with temperature-scaled similarity scoring; the mutual supervision strategy ensures all three modalities converge in the same latent space, enabling bidirectional translation between any pair. Precomputed embeddings span NCBI (700M sequences), GOPC (2B sequences), and OMG/MGnify (3B+ metagenomic sequences), indexed with FAISS for sub-second approximate nearest-neighbor retrieval at billion scale.

#Applications

ProTrek is particularly well suited to protein discovery workflows where the query is functional rather than sequence-based. Drug discovery teams can search for novel proteins sharing functional annotations with known targets using natural-language descriptions, bypassing the requirement for a seed sequence or structure. Structural biologists benefit from cross-modal retrieval — a newly resolved structure can be searched against functionally annotated databases directly. Metagenomic researchers gain access to precomputed embeddings for over three billion environmental protein sequences, enabling rapid functional annotation of proteins with no known homologs. The model's strong transfer performance also makes it a practical drop-in replacement for ESM-2 as a general-purpose protein encoder in supervised prediction pipelines.

#Impact

ProTrek's publication in Nature Biotechnology marks a significant step toward integrating all three primary axes of protein information — sequence, structure, and function — within a single retrieval framework. The publicly accessible search server and open model weights lower the barrier to adoption, particularly for experimental biologists who lack the infrastructure to run local homology searches. The model's demonstrated improvements over ESM-2 on downstream tasks suggest that multimodal pre-training provides richer protein representations than sequence-only approaches. A current limitation is that ProTrek encodes function through text annotations, meaning proteins with sparse or inaccurate database annotations may be poorly represented in the functional modality; retrieval quality therefore depends in part on the quality of the underlying protein databases.

Citation

ProTrek: Navigating the Protein Universe through Tri-Modal Contrastive Learning

Su, J., Zhou, X., Zhang, X., & Yuan, F. (2025). ProTrek: Navigating the Protein Universe through Tri-Modal Contrastive Learning. Nature Biotechnology.

DOI: 10.1038/s41587-025-02836-0

Recent citations

Papers that recently cited this model.

  • ProLoc: Text-guided Localization of Protein Functional Regions

    Peishuo Liu, Jiaxin Fan, Mianzhi Pan, et al.

    bioRxiv · Jul 2026

    0Influential
  • Advancing bioinformatics with language models: components, applications, and perspectives

    Jiajia Liu, Mengyuan Yang, Yankai Yu, et al.

    Briefings in Bioinformatics · Jul 2026

    0
  • Enzyme Kinetic Parameter Prediction via Catalytic Pocket-Augmented Machine Learning

    Ding Luo, Huining Ji, Shuming Cheng, et al.

    ACS Catalysis · Jun 2026

    0

Top citations

The most-cited papers that cite this model.

  • Leveraging Biomolecule and Natural Language through Multi-Modal Learning: A Survey

    Qizhi Pei, Lijun Wu, Kaiyuan Gao, et al.

    arXiv.org · Mar 2024

    27
  • Democratizing protein language model training, sharing and collaboration.

    Jin Su, Zhikai Li, Tianli Tao, et al.

    Nature Biotechnology · Oct 2025

    8
  • Deep learning and generative artificial intelligence methods in enzyme and cell engineering.

    Steffen Docter, Benoit David, Holger Gohlke

    Current Opinion in Biotechnology · Dec 2025

    4
  • Ab-initio amino acid sequence design from protein text description with ProtDAT

    Xiao-Yu Guo, Yi-Fan Li, Yuan Liu, et al.

    Nature Communications · Nov 2025

    2
  • “Visualize, Explore, and Select”: A Protein Language Model-based Approach Enabling Navigation of Protein Sequence Space for Enzyme Discovery and Mining

    Felix Moorhoff, David Medina-Ortiz, Alicja Kotnis, et al.

    bioRxiv · Mar 2026

    1

Related models

Models with similar goals, methods, or subject matter.

  • ProtST

    DeepGraphLearning

    Multi-modal protein language model trained on sequences paired with biomedical text, enabling zero-shot function prediction and text-based retrieval.

    Protein
  • ProtAlign

    Lawrence Livermore National Laboratory

    Cross-modal protein encoder that aligns ESM-2 sequence embeddings with ProteinMPNN structure embeddings in a shared space for cross-modal retrieval.

    Protein
  • ProCyon

    Harvard Medical School / Kempner Institute

    Multimodal foundation model integrating protein sequence, structure, and natural language to model and generate protein phenotypes across scales.

    ProteinLanguage modelSmall molecule
  • Unified Protein Embedding Model

    Harvard Medical School

    Siamese protein language model whose embedding distances approximate TM-score and lDDT, enabling alignment-free protein structure comparison.

    Protein
  • ProLoc

    Nanjing University

    Text-guided localization model that grounds natural-language functional descriptions to specific residue regions of a protein sequence.

    ProteinLanguage model

Citations

Total Citations24
Influential3
References39

GitHub

Stars210
Forks25
Open Issues7
Contributors3
Last Push2mo ago
LanguagePython
LicenseMIT

HuggingFace

Downloads44
Likes8
Last Modified11mo ago

Fields of citing research

  • Biology91%
  • Computer Science91%
  • Medicine48%
  • Chemistry17%
  • Engineering13%
  • Environmental Science4%
  • Physics4%

Share of papers citing this model.

Openness

bio.rodeo opennessOpen weights · open weights, closed recipe
66Partial
Usability — can I run it?100
Reproducibility — can I retrain it?18
open weights, closed recipe
Model Openness Framework
Class III
Open Model

Tags

contrastive_learningfoundation_modelmultimodalretrieval

Resources

GitHub RepositoryResearch PaperResearch PaperOfficial WebsiteHuggingFace ModelDataset