Managed notebooks for fine-tuning open protein, DNA and molecule models from labelled sequence data, with no GPU setup and linked cloud storage.
Protein language model family from 8M to 15B parameters, used as a frozen sequence encoder whose representations encode atomic-level structure.
Suite of six protein language models, including ProtBERT and ProtT5, trained on up to 393 billion amino acids without multiple sequence alignments.
Multi-species genomic foundation model swapping k-mer tokenization for byte pair encoding, matching Nucleotide Transformer with 21x fewer parameters.
DeepChem / Reverie Labs / Deep Forest Sciences / MIT CSAIL / UC Berkeley / University of Toronto
Released October 19, 2020
Chemical language model pretrained on up to 77 million PubChem SMILES strings for molecular property prediction on the MoleculeNet benchmarks.
DNA foundation models from 500M to 2.5B parameters, trained on 3,200+ human genomes and 850 species for variant effect prediction.
Sign-up is self-serve and requires no card. The free tier runs on spot capacity with a session time limit and a daily session cap. The notable design decision is storage: an existing S3, Cloud Storage or Blob bucket can be linked and worked against in place rather than uploaded, keeping data under the customer's own account and access policy. Enterprise single sign-on is reserved for the enterprise plan.
A complete rate card is published, down to overage rates. Four tiers run from free through a discounted researcher plan to professional and a negotiated enterprise tier, each bundling a monthly credit allowance. Credits meter per hour of active notebook use at a rate set by the machine, from shared CPU at the bottom to high-memory instances at the top. Two details shape planning: unused credits expire at cycle end, and linked storage is billed by the cloud provider.
Last verified