bio.rodeo
ModelsOrganizationsProvidersLeaderboardAboutSign in
bio.rodeo

The authoritative source for evaluating biological foundation models. No hype, just honest analysis.

Categories
  • DNA & Gene
  • RNA
  • Protein
  • Small molecule
  • Single-cell
  • Spatial omics
  • Pathology
  • Imaging
  • Metabolomics
  • Biosignals
  • Language model
bio.rodeoModelsOrganizationsProvidersLeaderboardAboutFAQSubmit a modelContact
© 2026 Pulsatance. All rights reserved. ~
Built by Pulsatance
models / protein / mp4-310-ai
ProteinLanguage model
310 AIUC BerkeleyMcGill UniversityReleased March 2025

MP4

Text-to-protein generative model designing de novo sequences from plain-language function descriptions, with designs confirmed by crystallography.

The short version

  • —Describe a function in plain English and get sequences back with no scaffold to start from
  • —Prompts combine organism, fitness criteria, and sequence hints in one specification
  • —70 synchronized training heads cover fold, functional site, and sequence novelty at once
  • —Designs expressed, crystallized to atomic resolution, and hydrolyzed ATP in vitro
10Openness0Citations

Where to run it

No providers recorded yet. Browse all providers

Most computational protein design starts from a scaffold. The designer picks a sequence, backbone, or motif that plausibly supports the target function, then optimizes around it — a workflow confined to whatever neighborhood of protein space the starting point occupies, and one that takes expertise to choose that starting point well.

MP4 — the Molecular Programming model, version 4 — from 310 AI removes that step. It is a transformer trained to map a natural-language prompt directly to an amino acid sequence, with no sequence or structural input required. A prompt can specify the desired function, the source organism, physical properties, and optional sequence hints; the model returns a sequence intended to satisfy them. The approach sits alongside ESM-3, Pinal, and ProteinDT in the text-conditioned design family, but MP4 consumes text alone at inference rather than routing through structure generation or controlled-tag conditioning.

What makes the preprint notable is how far the validation goes. Rather than stopping at in-silico metrics, the authors cloned 94 of 96 benchmark designs, expressed them, measured thermostability, solved two crystal structures at 1.30 Å and 1.77 Å, and demonstrated ATP binding and hydrolysis for designs prompted as ATPases. One of those structures has no close match in the Protein Data Bank.

#Key Features

  • Text-only conditioning: The model accepts a free-text prompt describing fitness criteria, source organism, and sequence properties, tokenizes it through a text-to-feature preprocessing stage, and generates a sequence with no template.
  • Multi-task output heads: Generation is routed through 70 synchronized task heads covering objectives such as structural fold determination, functional site prediction, and sequence novelty, whose outputs are combined into the final sequence.
  • Family-aware representation: Embedding natural proteins with MP4 and projecting with t-SNE separates them into distinct Pfam clan clusters, indicating the learned space organizes proteins by functional family.
  • Wet-lab validated designs: Of 94 clonable benchmark constructs, 79 (84%) produced detectable protein in a prokaryotic cell-free system, and 17 designs with reliable nanoDSF melting curves showed a mean apparent melting temperature above 62 °C, the most stable approaching 90 °C.
  • Demonstrated catalysis: Six ATP-binding designs — three ABC transporter nucleotide-binding domains and three adenylate kinases — showed ligand-induced thermal shifts of 2.4 °C to 7.5 °C, and three hydrolyzed ATP in ADP-Glo and Kinase-Glo assays.

#Technical Details

MP4 is a multi-layer transformer with multi-context sub-models that process different facets of the prompt features, an encoder aggregating them into a context-aware latent representation, a decoder, and the multi-task heads. Training used more than 3.2 billion datapoints from repositories including UniProt — 1.3 billion for sequences shorter than 300 residues and 2.3 billion for sequences shorter than 500 — over a unified 138,000-token vocabulary spanning natural-language descriptors and protein sequences, at roughly 3,800 AMD Instinct GPU-hours. Internal layer configurations, hyperparameters, and optimization strategies are proprietary and undisclosed.

The benchmark drew more than 1,000 prompts covering enzymatic activities, binding partners, and subcellular localization, from which 96 sequences were selected for novelty (under 50% identity to the NR database) and diversity. Against ESM-3 (fed InterPro classifications from the same prompts, without templates) and Pinal (fed the prompts as text), MP4 scored better on amino acid composition similarity to UniProt, on k=2 repetitiveness, on median ESMFold pLDDT, and on a recall-style overlap between prompt terms and predicted annotations. The comparison is not symmetric: Pinal builds structural templates internally, so it receives more starting information than MP4 or ESM-3, and the authors left that step enabled. Of the two crystallized designs, the 120-residue M1X0B overlays a known response-regulator fold at 0.938 Å RMSD, while the 77-residue homodimer MIYEI diverges from its closest deposited structures by 7.25 Å and 5.53 Å RMSD — a new fold.

#Applications

MP4 targets the earliest stage of a design campaign, where a researcher knows the function they want but not which scaffold to build on. Prompting for an activity, a ligand, or a localization returns candidate sequences that can go straight to synthesis and screening, lowering the expertise barrier for teams without a structural biology background and widening the sequence space explored by those who have one. The reported ATPase work illustrates the intended loop: prompt, filter computationally, express, and assay.

#Impact

The preprint has not been peer reviewed, and the authors are explicit that functional coverage, controllability, and interpretability remain open problems. Neither code nor weights are public — 310 AI cites commercial confidentiality and directs researchers to contact the company, with hosted inference offered through its browser-based 310 Copilot platform. What the work does establish is that a text-only interface can produce proteins that express, fold, crystallize, and catalyze, placing MP4's designed ATPases — three of them with directly measured hydrolysis — among the small canon of machine-learning-designed enzymes that includes pGAN59, FastPETase, ProGen's L056, and the PLACER-designed serine hydrolase momi120_102.

At a glance

Released
March 2025
Category
Protein
Organizations
310 AI / UC Berkeley / McGill University

Related models

  • ProDVa

    East China Normal University / TeleAI / Fudan University

  • InstructPro

    Carnegie Mellon University / Lambda

  • ProteinMPNN

    Institute for Protein Design

  • Pinal

    Westlake University

  • ProteinDT

    UC Berkeley

  • Sparks

    MIT

Links

bioRxiv PreprintOfficial WebsiteDocumentation

Tags

de_novo_designenzyme_designgenerativemulti_taskprotein_designproteomicstransformer

Something wrong?

Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.