tf-SFM

Transcription factor-DNA binding specificity prediction from sequence, with a physics-derived dual-encoder trained by symmetric contrastive learning.

Released: June 2026

tf-SFM is a Specificity Foundation Model (SFM) for predicting transcription factor (TF)–DNA binding specificity directly from sequence. Determining which TF recognizes which genomic sequence is fundamental to understanding gene regulation, and conventional approaches rely on experimental assays such as protein binding microarrays, SELEX, or ChIP-seq. tf-SFM instead frames TF–DNA recognition as a cross-modal matching problem, learning to align cognate protein–DNA pairs in a shared representation space so that likely binding partners can be scored and retrieved from sequence alone.

Developed by the Reddy lab at ETH Zurich and posted as a bioRxiv preprint in June 2026, tf-SFM is one of six models in the SFM family, all built on a single, physics-derived dual-encoder architecture. It is the sequel to CALM-1, the antibody–antigen specificity model from the same group, generalizing that contrastive molecular-recognition recipe from immune binding to gene-regulatory binding.

The model encodes TF and DNA sequences with separate encoders and aligns them using a symmetric contrastive objective, pulling true binding pairs together and pushing non-binders apart. This formulation lets tf-SFM transfer knowledge learned across many regulatory contexts into zero-shot predictions on held-out TFs and binding sites.

Key Features

Sequence-to-specificity prediction: Predicts TF–DNA binding from protein and nucleotide sequence alone, without requiring structural data or motif models.
Physics-derived dual-encoder: Encodes transcription factor and DNA separately, with an architecture motivated by the physics of molecular recognition rather than a generic backbone.
Symmetric contrastive learning: Aligns cognate TF–DNA pairs in a shared embedding space, enabling retrieval in either direction (TF-to-site or site-to-TF).
Learned Boltzmann temperature: A learned temperature parameter calibrates the contrastive similarity scores in a thermodynamically motivated way.
Zero-shot cross-modal retrieval: Generalizes to unseen TFs and binding sites without task-specific fine-tuning.

Technical Details

tf-SFM uses the shared SFM architecture: a physics-derived dual-encoder trained with a symmetric contrastive objective and a learned Boltzmann temperature that calibrates similarity scores. The two encoders embed transcription factor and DNA sequences independently, and the contrastive loss aligns cognate pairs while separating mismatches. The model is pretrained on public TF–DNA specificity data and evaluated by zero-shot cross-modal retrieval on held-out pairs, where it reports strong top-k retrieval performance—mirroring the benchmarks used across the SFM family for measuring how reliably a model recovers true binding partners.

Applications

tf-SFM is aimed at regulatory genomics, where identifying or prioritizing TF–DNA interactions from sequence can accelerate the annotation of regulatory elements and the design of synthetic promoters. By scoring and retrieving likely binding partners, it can help generate hypotheses about which factors drive a given regulatory site, triage candidate binding sites for a TF of interest, and complement experimental assays in settings where wet-lab characterization is limited.

Impact

tf-SFM extends the contrastive specificity-prediction paradigm established by CALM-1 from antibody–antigen recognition to transcription factor–DNA recognition, demonstrating that a single physics-derived dual-encoder recipe transfers across molecular domains. As one of six SFMs released together, it contributes evidence that cross-modal contrastive learning is a general tool for biological specificity prediction. Its main current limitations are those of a recent preprint: results await peer review and independent benchmarking, and at the time of release no public code or weights repository was available, so reproduction depends on forthcoming artifact releases.

Citation

Vibe Coding Specificity Foundation Models

Reddy, S. T. (2026) Vibe Coding Specificity Foundation Models. bioRxiv.

DOI: 10.64898/2026.06.04.730134

Recent citations

Papers that recently cited this model.

Generative Drug Design in a Loop with dtSFM
Sai T. Reddy
bioRxiv · Jun 2026
0
A Drug–Target Specificity Foundation Model for Off-target Prediction, Repurposing, and Generative Design
Sai T. Reddy
bioRxiv · Jun 2026
0

Top citations

The most-cited papers that cite this model.

A Drug–Target Specificity Foundation Model for Off-target Prediction, Repurposing, and Generative Design
Sai T. Reddy
bioRxiv · Jun 2026
0
Generative Drug Design in a Loop with dtSFM
Sai T. Reddy
bioRxiv · Jun 2026
0

Citations

Total Citations2

Influential0

References52

Fields of citing research

Biology100%
Computer Science100%
Medicine100%
Chemistry50%

Share of papers citing this model.

Openness

bio.rodeo opennessClosed · low usability and reproducibility

18Closed

Usability — can I run it?14

Reproducibility — can I retrain it?12

Model Openness Framework

Unclassified

Missing required components

Resources

Research Paper

Key Features

Sequence-to-specificity prediction: Predicts TF–DNA binding from protein and nucleotide sequence alone, without requiring structural data or motif models.

Physics-derived dual-encoder: Encodes transcription factor and DNA separately, with an architecture motivated by the physics of molecular recognition rather than a generic backbone.

Symmetric contrastive learning: Aligns cognate TF–DNA pairs in a shared embedding space, enabling retrieval in either direction (TF-to-site or site-to-TF).

Learned Boltzmann temperature: A learned temperature parameter calibrates the contrastive similarity scores in a thermodynamically motivated way.

Zero-shot cross-modal retrieval: Generalizes to unseen TFs and binding sites without task-specific fine-tuning.

Technical Details

Applications

Impact

tf-SFM

Key Features

Technical Details

Applications

Impact

Citation

Vibe Coding Specificity Foundation Models

Recent citations

Generative Drug Design in a Loop with dtSFM

A Drug–Target Specificity Foundation Model for Off-target Prediction, Repurposing, and Generative Design

Top citations

A Drug–Target Specificity Foundation Model for Off-target Prediction, Repurposing, and Generative Design

Generative Drug Design in a Loop with dtSFM

Citations

Fields of citing research

Openness

Tags

Resources

tf-SFM

Key Features

Technical Details

Applications

Impact

Citation

Vibe Coding Specificity Foundation Models

Recent citations

Generative Drug Design in a Loop with dtSFM

A Drug–Target Specificity Foundation Model for Off-target Prediction, Repurposing, and Generative Design

Top citations

A Drug–Target Specificity Foundation Model for Off-target Prediction, Repurposing, and Generative Design

Generative Drug Design in a Loop with dtSFM

Citations

Fields of citing research

Openness

Tags

Resources

tf-SFM

#Key Features

#Technical Details

#Applications

#Impact

Citation

Vibe Coding Specificity Foundation Models

Recent citations

Top citations

Related models

Citations

Fields of citing research

Openness

Tags

Resources

tf-SFM

#Key Features

#Technical Details

#Applications

#Impact

Citation

Vibe Coding Specificity Foundation Models

Recent citations

Top citations

Related models

Citations

Fields of citing research

Openness

Tags

Resources

Key Features

Technical Details

Applications

Impact

Key Features

Technical Details

Applications

Impact