Open-source all-atom co-folding foundation model for protein-ligand, protein-protein, and antibody-antigen complex prediction in drug discovery.
Autoregressive language model trained on 37 million intrinsically disordered region sequences, generating IDRs given flanking folded domains.
Structure-based drug design model that unifies de novo generation, docking, conformer generation, and pharmacophore conditioning via flow matching.
Dual-target protein sequence design conditioned on two receptor structures at once, combining a heterogeneous graph network with ESM-2 features.
Sequence-based multitask model predicting covalently ligandable cysteines and reversible ligand-binding residues across the human proteome.
Protein sequence encoder that maps ESM2 embeddings to a learned 20-letter alphabet for structure-quality remote homology detection at MMseqs2 speed.
Protein interface prediction from sequence alone, swapping hand-crafted features for frozen ProtT5-XL embeddings that hold up on remote homologs.
Antibody and TCR CDR sequence design by structure retrieval, matching query loops against solved CDR structures rather than generating residues.
Binding energy for protein-ligand, protein-protein, and antibody-antigen complexes is read off an energy model trained on crystal structures alone.
Autoregressive 3D structure model built on an octree tokenizer, spanning molecule generation, molecular docking, and protein pocket prediction.
Cryo-EM density-map-to-atomic-structure modeling that fuses protein language model embeddings with density voxels, then refines with AlphaFold3.
Inverse protein folding from backbone coordinates, chaining a pretrained structure encoder into a pretrained sequence autoencoder on small data.
Sequence-based binding site predictor spanning protein-DNA, protein-RNA, protein-protein, and antibody-antigen interfaces via a fine-tuned ProtT5.
Dual-encoder contrastive model that retrieves enzymes for query reactions by matching reaction fingerprints to protein sequence embeddings.
Generative protein-dynamics model that predicts short molecular dynamics trajectories with rectified flow matching over residue frames and torsions.
Kinase-substrate specificity prediction from sequence alone, using ESM-2 embeddings to score phosphorylation across whole mammalian kinomes.
Protein conformational ensemble generator that samples heavy-atom structures in a latent space, with a variant conditioned on temperature.
All-atom structure prediction for complexes of proteins, DNA, RNA, and small molecules, using Min-SNR diffusion weighting and the Muon optimizer.
RNA language model pre-trained on 2M+ pre-mRNA sequences from 72 vertebrate species for splice-site prediction and variant effect analysis.
RNA language model pretrained on 30M non-coding RNA sequences that predicts secondary structure, contacts, and splice sites without alignments.
Protein language model trained from scratch on MD and normal-mode dynamics, representing residue fluctuation and co-movement from sequence.
Predicts protein complex stoichiometry from amino acid sequence alone, ranking copy numbers in seconds and exporting AlphaFold3-ready JSON files.
Sequence-based protein-ligand binding site predictor pairing a protein language model with a SMILES chemical language model for zero-shot ligands.