
Metabolite profiles from NMR and mass spectrometry
20 models in this category
Metabolomics foundation models learn from metabolite profiles measured by nuclear magnetic resonance spectroscopy and mass spectrometry, capturing the small-molecule signatures of cellular metabolism across thousands of detected features. Unlike genomics or transcriptomics, metabolomics data reflects the downstream chemistry of both genetics and environment, making it particularly informative for physiological phenotyping. These models aim to learn representations that generalize across instruments, sample types, and experimental conditions — challenges that have historically fragmented the metabolomics field and limited cross-study comparability.
Biomarker discovery — identifying metabolites or metabolite signatures that distinguish disease states, treatment responses, or physiological conditions — is the primary application driving interest in foundation models for metabolomics. Metabolic phenotyping of population cohorts, where models must integrate hundreds to thousands of measured features across diverse participants, benefits from learned representations that capture co-regulation patterns invisible to univariate approaches. Spectral annotation, the task of identifying unknown peaks from NMR or MS spectra, is another area where pretrained models have begun to reduce the manual curation bottleneck.
Top-rated metabolomics models from our evaluations
Tandem mass spectrometry model that embeds MS/MS spectra and molecular graphs in one space, ranking candidate structures without a spectral library.
Self-supervised transformer pretrained on millions of tandem mass spectra, giving embeddings for spectral annotation and fingerprint prediction.
Multimodal conversational LLM for metabolite analysis, fusing a molecular-graph GNN and molecular-image CNN with a Vicuna-13B language backbone.
Vision Transformer foundation model for spatial metabolomics, pretrained on ~4,000 curated METASPACE mass spectrometry imaging datasets.
Foundation model for tandem mass spectrometry that embeds MS/MS spectra into a learned chemical space, resolving isomers and classifying disease.
Contrastive encoder aligning NMR metabolomics to the plasma proteome, adding proteome-level disease risk signal to cohorts with no proteomics.
A metabolomics foundation model is a neural network pretrained on large collections of metabolomic measurements — mass spectrometry or NMR profiles capturing hundreds to thousands of metabolite features per sample — to learn representations of metabolic state that transfer to tasks like biomarker discovery and phenotype prediction. The field is earlier-stage than genomics or proteomics foundation modeling, but the scale of population metabolomics datasets is creating conditions for pretraining at meaningful scale.
Metabolomics datasets are highly heterogeneous: different instruments, ionization methods, and sample preparation protocols produce feature spaces that don't align without careful harmonization. Missing values are common because many metabolites fall below detection thresholds, and the total number of reliably detected metabolites per study is much smaller than the gene count in transcriptomics, limiting the dimensionality available for representation learning. Cross-study generalization requires models that are robust to these technical variations.
Metabolomics captures the downstream chemistry of gene expression and environmental exposure, reflecting both genetic regulation and lifestyle factors like diet, microbiome activity, and drug metabolism that are poorly represented in genomic or transcriptomic data alone. Multi-omics approaches that combine metabolomics with genomics or transcriptomics can therefore capture complementary axes of biological variation. Foundation models that integrate across these data types — multi-omics models — represent an active research direction.