Protein structure and complex prediction from sequence, in a three-track network that reasons over alignments, distances, and 3D coordinates at once.
No providers recorded yet. Browse all providers
A two-track folding network reasons about sequence and about a residue-residue distance map together, then hands the finished distance map to a structure module that builds coordinates. The coordinates never talk back. RoseTTAFold's premise is that they should: a third track operating directly on 3D backbone coordinates runs in parallel with the 1D alignment track and the 2D distance track, information flowing in both directions among all three throughout the network. A sequence position can be revised because of where that residue ended up in space, not only the other way around.
The architecture came out of the Baker Lab at the University of Washington Institute for Protein Design and was published in Science in July 2021. Its context was CASP14, where AlphaFold 2 had demonstrated an accuracy jump but had not published its method. Baek and colleagues worked from the five architectural ideas described at the conference, built a two-track network that already beat trRosetta, then extended it to three tracks — and released trained weights and running code roughly two weeks before AlphaFold 2's own code appeared.
Both networks ship from the same repository. The three-track model reaches a structure two ways: a pyRosetta route that builds relaxed all-atom models from predicted distance and orientation distributions on an 8 GB GPU, and an end-to-end route that emits backbone coordinates through a final SE(3)-equivariant layer. The two-track network followed in November 2021 as RoseTTAFold-2track, or RF2t: a 10.7-million-parameter checkpoint with no coordinate track, predicting inter-chain contacts and nothing else.
The 1D track uses Performer-style attention over the multiple sequence alignment, the 2D track attends over the distance map, and the 3D track is built from SE(3)-equivariant transformer layers. Training used discontinuous crops of two sequence segments totalling 260 residues, averaged across crops at inference, because GPU memory ruled out whole proteins. Alignments come from HHblits against UniRef30 and the 272 GB BFD, templates from HHsearch against pdb100, capped at the top 1000 sequences. After roughly 1.5 hours of that search, the end-to-end version folds a sub-400-residue protein in about 10 minutes on an RTX 2080, a complex in 30 minutes on a 24 GB TITAN RTX. RF2t needs 11 seconds for a 1,000-residue paired alignment on the same card, about 100 times faster than AlphaFold 2.
On CASP14 targets the three-track model beat both leading server groups and the second-ranked human group while remaining short of AlphaFold 2, and it topped every server evaluated on CAMEO's medium and hard targets. Over a third of its models for 693 human domains without close structural homologs cleared predicted lDDT 0.8.
The original paper's strongest evidence is experimental rescue. Four crystallographic datasets that had resisted molecular replacement with every available PDB model — among them a bacterial surface layer protein and the fungal protein Lrbp — were phased with RoseTTAFold models where trRosetta had failed. Elsewhere the same predictions fit a low-resolution PI3Kγ cryo-EM map and placed pathogenic mutations in TANGO2, the ADAM33 prodomain and ceramide synthase CERS1 against predicted active sites.
RF2t took the network somewhere the three-track model could not afford to go. Humphreys and colleagues scored all 4.3 million paired alignments buildable from 8.3 million yeast protein pairs, separating gold-standard complexes from random pairs far better than direct coupling analysis. AlphaFold 2, too slow for a sweep that size, then modelled the 5,495 highest-scoring pairs. Together the two identified 1,505 likely interactions and produced structures for 106 previously unidentified assemblies and 806 complexes never structurally characterised, all deposited in ModelArchive.
RoseTTAFold made high-accuracy folding available to any lab with a gaming GPU at a moment when it was otherwise unobtainable, and the Robetta server to labs with none. The code is MIT-licensed; the trained weights, RF2t.pt included, carry the Rosetta-DL agreement, non-commercial research only — a split that still governs much of the family. RFdiffusion turns the network into a generative denoiser for de novo design, RoseTTAFold2 rebuilds it with AlphaFold 2 training ideas folded in, RoseTTAFold All-Atom extends it to nucleic acids and small molecules, and RF3 carries the lineage into all-atom diffusion. The limits stay plain too: no training code was ever released, and inference still waits on a slow alignment search.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.